9 days ago
Pune, IndiaSenior
Responsibilities
- Define and refine SLIs, SLOs, and SLAs for critical services with product and engineering leadership.
- Own error budget tracking and policies and drive reliability investments when budgets are at risk.
- Design and implement observability platforms using metrics, structured logging, and distributed tracing.
- Automate repetitive operational work and lead measurable toil-reduction initiatives.
- Design and execute chaos engineering programs to identify reliability weaknesses proactively.
- Facilitate blameless postmortems and track corrective actions to completion.
- Build incident-response processes, runbooks, escalation paths, and healthy on-call rotations.
- Design, build, and support resilient, observable, secure, and cost-optimized cloud infrastructure.
- Automate production deployments on AWS and improve CI/CD pipelines with reliability and quality gates.
- Develop fault-tolerant management solutions across multiple cloud platforms and data centers.
- Conduct production-readiness reviews and launch-checklist activities for new services and features.
- Champion SRE practices, mentor engineers, participate in roadmap planning, and evaluate technologies.
Requirements
- At least 8 years of experience with scalable, distributed-systems architecture.
- At least 3 years of hands-on Site Reliability Engineering experience, including ownership of SLOs and error budgets.
- At least 4 years of cloud-platform experience, including AWS.
- At least 4 years of infrastructure-as-code experience with Terraform and AWS CDK.
- At least 5 years of scripting experience using Python, Shell, or a similar language.
- At least 3 years of containerization experience, including Docker.
- At least 4 years of orchestration experience, including Kubernetes.
- Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, OpenTelemetry, and Jaeger.
- Experience with incident-management platforms and on-call tooling such as PagerDuty and OpsGenie.
- Experience implementing automated service deployments covering networking, security, reliability, management, reporting, and configuration management.
- Experience with chaos-engineering principles and tools such as Chaos Monkey, LitmusChaos, and Gremlin.
- Experience managing PostgreSQL, Redis, DynamoDB, and MongoDB databases.
- Strong understanding of deployment automation, production-readiness reviews, networking, HTTP, and distributed-system latency diagnosis.
- Experience using Git in a team environment and working with distributed tracing frameworks such as Jaeger, Zipkin, and AWS X-Ray.
- Preferred experience includes Google SRE principles, AIOps or ML-based anomaly detection, capacity planning, load testing, performance benchmarking, CI/CD reliability gates, and customer-facing production issue resolution.
- CS degree or equivalent experience.
Benefits
- Hybrid work arrangement in Pune.
- Equal opportunity employment and commitment to a diverse workplace.
Tech Stack
Categories
DevOpsSite Reliability
