Responsibilities
- Design and implement scalable, fault-tolerant batch and streaming data pipelines.
- Lead the design and development of reusable data platforms and frameworks for multiple teams and use cases.
- Build and optimize data models and schemas for operational and analytical workloads.
- Understand, customize, and extend Apache Spark internals and open-source code.
- Develop streaming solutions using Apache Flink and Spark Structured Streaming.
- Abstract infrastructure complexity to enable ML, analytics, and product teams to build more efficiently.
- Promote platform reusability, extensibility, and developer self-service.
- Ensure data quality, consistency, governance, observability, and appropriate access controls.
- Optimize cloud-native infrastructure for cost, latency, performance, and scalability.
- Mentor junior engineers, participate in architecture reviews, and uphold engineering standards.
- Collaborate with product, ML, and data teams on technical solutions.
Requirements
- 5–8 years of professional software or data engineering experience focused on distributed data systems.
- Strong programming skills in Java, Scala, or Python and expertise in SQL.
- At least 2 years of hands-on experience with Apache Kafka, Apache Spark, EMR or Dataproc, Hive, Delta Lake, Presto or Trino, Airflow, and data lineage tools.
- Experience tuning Spark, Delta Lake, and Presto at terabyte scale or beyond.
- Strong understanding of Spark internals, including Catalyst, Tungsten, and shuffle, with experience contributing to or customizing open-source code.
- Familiarity with Apache Iceberg, Hudi, Trino or Presto, DuckDB, ClickHouse, Pinot, Druid, Airflow, Dagster, Prefect, DBT, Great Expectations, DataHub, OpenMetadata, Kubernetes, Terraform, and Docker.
- Exposure to data security, privacy, observability, and compliance frameworks is preferred.
- Open-source contributions in the big data ecosystem are preferred.
- Experience with data modeling and end-to-end data pipeline development is preferred.
- Familiarity with OLAP data cubes and Tableau, Power BI, Superset, or Looker is preferred.
- Working knowledge of ELK Stack, Redis, MySQL, RxJava, Spring Boot, and microservices is preferred.
Benefits
- Competitive cash and equity-based compensation tailored to role, experience, and skills.
- Extensive medical insurance for employees and families, telehealth, wellness events, and fitness-related perks.
- Generous leave policies, parental support, retirement benefits, and learning and development assistance.
- Recognition programs, employee activities, salary advance support, relocation assistance, and flexible benefit plans.
- Inclusive and accessible workplace with reasonable interview accommodations and accessibility support.
Tech Stack
Categories
About Meesho
Meesho is India’s e-commerce marketplace, on a mission to democratise internet commerce. Our multi-sided technology platform connects four key stakeholders — consumers, sellers, logistics partners, and content creators — to power inclusive growth at scale. We enable individuals and small businesses to sell online with ease, offering access to a wide customer base, integrated logistics, payment solutions, and platform support. For customers, Meesho offers a broad and affordable selection, tailored for diverse needs across Bharat. We also empower creators to build commerce-driven content that drives discovery and engagement. Our logistics operations are powered by Valmo, Meesho’s asset-light logistics platform that works entirely through partner-led infrastructure to ensure cost-efficient and scalable deliveries.
