21 days ago
Noida, IndiaMid Level / Senior
Responsibilities
- Design, build, maintain, and optimize high-volume ETL/ELT pipelines across AWS and the Hadoop ecosystem.
- Develop distributed data processing solutions using PySpark, Spark SQL, Scala, and serverless cloud patterns.
- Implement reusable ingestion frameworks and orchestrate batch and event-driven workflows using Step Functions, Airflow, Control-M, or similar tools.
- Optimize data workflows through partitioning, bucketing, compression, and Parquet or ORC file formats.
- Design hybrid data lake architectures using Amazon S3 and HDFS while supporting data governance, security, and compliance.
- Define technical specifications, make scalable architecture decisions, and implement secure, maintainable software solutions.
- Monitor and troubleshoot incidents, Spark performance issues, job failures, cluster bottlenecks, and business-as-usual activities.
- Collaborate with business stakeholders, QA, and cross-functional teams to deliver high-quality datasets and streamline releases.
- Provide technical guidance and best-practice adoption support to peers and junior team members.
Requirements
- 4 to 8 years of overall ETL experience.
- Strong experience with Big Data technologies, including Spark and Cloudera.
- Hands-on expertise with Scala, PySpark, Spark optimization, HiveQL, and distributed computing.
- Strong experience with AWS data services and the AWS data stack.
- Good understanding of Hadoop, HDFS, Hive, Spark, YARN, Kafka, Hive SQL, and Impala.
- Proficiency in Python or Shell scripting and strong SQL experience.
- Experience with ETL, data warehousing, data modeling, star and snowflake schemas, partitioning strategies, and schema evolution.
- Experience with Airflow, Control-M, Step Functions, or other workflow orchestrators.
- Experience with GitHub, Git commands, and CI/CD practices.
- Familiarity with serverless patterns, Docker, ECS, and EKS.
- Understanding of data profiling, data flow diagrams, data lineage, end-to-end data architecture, and cloud security and compliance.
- Strong logical, analytical, problem-solving, communication, collaboration, and stakeholder coordination skills.
- AWS Data Engineer or Developer certification is a plus.
Benefits
- Health benefits are available starting Day One, along with wellbeing programs, retirement plans, and career growth opportunities.
- The company offers continuing education and training and potential for growth in a global organization.
- Virtual interviews are conducted by video, the company maintains a camera-on culture, and occasional travel to a physical office may be required.
Tech Stack
Amazon RedshiftApache AirflowApache HadoopApache HiveApache KafkaApache SparkAWSDockerGitPythonScalaYarn
Categories
Data Engineering
