Yotta Infrastructure

Data Platform Engineer

Yotta Infrastructure
Apply
3 months ago
Delhi, IndiaEntry Level

Responsibilities

  • Assist with setup, monitoring, maintenance, and troubleshooting of Apache Spark clusters, Apache Hive, Hadoop, and Airflow environments.
  • Support development and scheduling of data pipelines using Airflow DAGs and Python scripts.
  • Help configure and manage JupyterHub for multi-user access and Spark integration.
  • Monitor cluster health and performance and assist with troubleshooting Spark job failures.
  • Write and maintain Python automation scripts for data workflows, ETL, and process automation.
  • Participate in code reviews, documentation, and deployment activities.
  • Collaborate with senior engineers and data scientists to implement improvements and new features.

Requirements

  • At least 1 year of total or relevant experience.
  • Bachelor’s degree or another relevant degree.
  • Basic understanding of Apache Spark, Apache Hive, and Hadoop File System.
  • Familiarity with Apache Airflow, including DAGs, scheduling, and task dependencies.
  • Hands-on Python scripting experience for data processing, automation, or API interaction.
  • Comfort working with Linux and container environments, including command-line tools, system logs, and process management.
  • Understanding of ETL and distributed computing concepts.
  • Basic knowledge of Git and version control.
  • Preferred: exposure to Jupyter or JupyterHub, Docker or Kubernetes, SQL, structured and unstructured data, and cloud platforms such as AWS, GCP, or Azure.

Benefits

  • General shift schedule.
  • Three interview rounds.

Categories

Data EngineeringDevOps
Contact me