8 days ago
Bengaluru, IndiaSenior
Responsibilities
- Design and build scalable batch and real-time data pipelines using Spark, Flink, Kafka, and Airflow.
- Develop reliable ETL/ELT frameworks and data products supporting analytics, experimentation, recommendations, personalization, and AI applications.
- Build identity resolution, cross-surface signal aggregation, audience, user profiling, and personalization systems.
- Develop ingestion, enrichment, normalization, deduplication, taxonomy, and quality systems for commerce catalogs and merchant feeds.
- Build feature-generation frameworks and low-latency pipelines supporting ML training and online inference.
- Develop AI-powered internal tools for pipeline debugging, data-quality triage, SQL generation, metadata discovery, schema analysis, and cost optimization.
- Own production services and pipelines, including SLAs, observability, lineage, alerting, reconciliation, incident response, and root-cause analysis.
- Lead architecture and design discussions, establish engineering best practices, and mentor junior engineers.
Requirements
- 6–10 years of experience in data engineering, distributed systems, or data platform development.
- Strong experience owning large-scale production systems end to end.
- Hands-on experience with Apache Spark, Kafka, Flink, Airflow, distributed data processing, and batch and streaming architectures.
- Strong understanding of dimensional modeling, data warehousing, large-scale schema design, and evolving schemas.
- Experience with data validation, lineage, monitoring and alerting, reconciliation pipelines, and data-system delivery practices.
- Experience with GCP, Databricks, BigQuery, infrastructure as code, cluster management, performance tuning, and cost optimization.
- Strong programming skills in Python and Scala or Java, plus SQL.
- Strong understanding of system design, distributed systems, performance optimization, and reliability engineering.
- Preferred experience with product catalogs, affiliate commerce, merchant feeds, search and recommendation systems, identity resolution, audience platforms, customer 360 systems, feature stores, training-data pipelines, real-time inference systems, MLOps infrastructure, and LLM-powered developer tools.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkDatabricksGoogle BigQueryGoogle Cloud PlatformJavaPythonScalaSQL
Categories
Data Engineering
