21 days ago
Noida, IndiaMid Level / Senior
Responsibilities
- Design, construct, test, and maintain scalable ETL/ELT data pipelines using Python, Spark, and PySpark.
- Architect and optimize distributed data processing systems using Apache Spark and Hadoop.
- Build and manage serverless and cloud-native data architectures on AWS, including S3, EMR, Glue, Redshift, Lambda, and Athena.
- Design dimensional data models and optimize query performance for large-scale analytics.
- Implement automated data quality checks, monitoring and alerting, and data governance and security standards.
Requirements
- Require 4–6 years of professional experience in data engineering or software development.
- Require hands-on experience building data lakes and data warehouses natively on AWS.
- Require proficiency with distributed computing engines such as Apache Spark and streaming technologies such as Kafka or Kinesis.
- Require strong Python programming skills and advanced SQL expertise.
- Require a solid understanding of data warehousing concepts and dimensional modeling.
- Prefer experience with Terraform or CloudFormation, Apache Airflow or Step Functions, Databricks, or Snowflake.
Benefits
- Health, dental, and vision coverage begins Day One, with wellbeing programs, retirement contribution matching, generous time off, parental leave, continuing education, and career growth opportunities.
- Flexible working arrangements are considered wherever possible, with occasional travel to a physical office potentially required.
- Virtual interviews are conducted by video and may require the applicant to be on camera.
Tech Stack
Amazon RedshiftApache AirflowApache HadoopApache KafkaApache SparkAWSDatabricksPythonSnowflakeSQLTerraform
Categories
Data Engineering
