
Data Platform Engineer
Yotta Infrastructure3 months ago
Delhi, IndiaEntry Level
Responsibilities
- Assist with setup, monitoring, maintenance, and troubleshooting of Apache Spark clusters, Apache Hive, Hadoop, and Airflow environments.
- Support development and scheduling of data pipelines using Airflow DAGs and Python scripts.
- Help configure and manage JupyterHub for multi-user access and Spark integration.
- Monitor cluster health and performance and assist with troubleshooting Spark job failures.
- Write and maintain Python automation scripts for data workflows, ETL, and process automation.
- Participate in code reviews, documentation, and deployment activities.
- Collaborate with senior engineers and data scientists to implement improvements and new features.
Requirements
- At least 1 year of total or relevant experience.
- Bachelor’s degree or another relevant degree.
- Basic understanding of Apache Spark, Apache Hive, and Hadoop File System.
- Familiarity with Apache Airflow, including DAGs, scheduling, and task dependencies.
- Hands-on Python scripting experience for data processing, automation, or API interaction.
- Comfort working with Linux and container environments, including command-line tools, system logs, and process management.
- Understanding of ETL and distributed computing concepts.
- Basic knowledge of Git and version control.
- Preferred: exposure to Jupyter or JupyterHub, Docker or Kubernetes, SQL, structured and unstructured data, and cloud platforms such as AWS, GCP, or Azure.
Benefits
- General shift schedule.
- Three interview rounds.
Tech Stack
Apache AirflowApache HadoopApache HiveApache SparkAWSAzureDockerGitGoogle Cloud PlatformKubernetesLinuxPythonSQL
Categories
Data EngineeringDevOps