8 days ago
Bengaluru, IndiaSenior / Staff+
Responsibilities
- Design and build scalable batch and real-time data pipelines using Spark, Flink, Kafka, and Airflow.
- Develop reliable ETL/ELT frameworks and data products supporting analytics, experimentation, recommendations, personalization, and AI applications.
- Build user identity resolution, cross-surface signal aggregation, audience, profiling, segmentation, and personalization systems.
- Develop commerce catalog ingestion, enrichment, normalization, deduplication, taxonomy, and quality systems.
- Build feature generation frameworks and low-latency pipelines for ML training and online inference.
- Develop AI-powered internal tools and self-service capabilities for pipeline debugging, data quality triage, SQL optimization, metadata discovery, schema analysis, and cost optimization.
- Own production pipelines and services, including SLAs, monitoring, lineage, alerting, reconciliation, quality checks, incident response, and root-cause analysis.
- Lead architecture and design discussions, drive engineering best practices, mentor junior engineers, and contribute to technical leadership.
Requirements
- 6–10 years of experience in Data Engineering, Distributed Systems, or Data Platform development.
- Strong experience owning large-scale production systems end-to-end.
- Hands-on experience with Apache Spark, Kafka, Flink, Airflow, distributed data processing, and batch and streaming architectures.
- Strong understanding of dimensional modeling, data warehousing, large-scale schema design, complex datasets, and evolving schemas.
- Experience with data validation frameworks, lineage systems, monitoring and alerting, reconciliation pipelines, and CI/CD for data systems.
- Experience with GCP, Databricks, BigQuery, infrastructure as code, cluster management, performance tuning, and cost optimization.
- Strong programming skills in Python, Scala or Java, and SQL.
- Strong understanding of system design, distributed systems, performance optimization, and reliability engineering.
- Preferred experience with product catalogs, affiliate commerce platforms, merchant feeds, search and recommendation systems, identity resolution, audience platforms, Customer 360 systems, user profiling, segmentation, feature stores, training data pipelines, real-time inference data systems, MLOps infrastructure, and LLM-powered developer tools.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkDatabricksGoogle BigQueryGoogle Cloud PlatformJavaPythonScalaSQL
Categories
Data Engineering
