8 days ago
Bengaluru, IndiaSenior
Responsibilities
- Own critical backend components and services end to end, including architecture, development, deployment, operations, and evolution.
- Create detailed low-level designs covering data models, APIs, concurrency patterns, failure modes, and system behavior.
- Write production-grade backend code daily and build systems operating at 200K+ queries per second and petabyte scale.
- Manage infrastructure, capacity planning, scaling, reliability improvements, monitoring, alerting, debugging, and cost optimization.
- Lead incident response, post-mortems, on-call operations, and systemic reliability improvements.
- Define technical roadmaps balancing feature delivery, technical debt, and operational improvements.
- Mentor junior and mid-level engineers on system design, code quality, and production practices.
- Collaborate with product, infrastructure, and engineering teams to define requirements and deliver solutions.
Requirements
- 4-5+ years of experience building and operating backend systems in production at scale.
- B.E./B.Tech in Computer Science or equivalent practical experience.
- Demonstrated end-to-end ownership of significant components or services from inception to maturity.
- Strong low-level design skills, including data structures, algorithms, API contracts, concurrency models, and failure handling.
- Expert proficiency in at least one modern programming language such as Go, Python, or Java.
- Deep hands-on experience with SQL and NoSQL databases, including schema design, query optimization, indexing, and operational troubleshooting.
- Extensive experience building, deploying, and operating high-throughput microservices and large-scale data pipelines.
- Expertise in observability, including metrics, logging, tracing, alerting, and production debugging.
- Significant experience with incident response, on-call rotations, live issue debugging, infrastructure management, capacity planning, and cost optimization.
- Strong distributed-systems fundamentals covering concurrency, consistency models, and failure scenarios.
- Preferred experience with GCP, AWS, or Azure; Kubernetes; Terraform or Pulumi; Kafka, Pub/Sub, Kinesis, or Flink; SRE and reliability practices; FinOps; performance optimization; technical mentorship; or open-source and technical-writing contributions.
