17 days ago
Bengaluru, IndiaSenior
Responsibilities
- Design and implement resilience strategies across AWS, Azure, and GCP environments.
- Support chaos engineering initiatives and integrate experiments into CI/CD pipelines.
- Validate RTO, RPO, SLO, and recovery objectives using metrics and governance controls.
- Improve resilience across databases, caching, messaging, and authentication services.
- Validate multi-region and cross-cloud failover, retry logic, circuit breakers, and graceful degradation.
- Use Grafana, Prometheus, and LogDNA for observability-driven validation.
- Integrate reliability checks into Jenkins pipelines and ensure production readiness before releases.
- Collaborate with SRE, platform, and development teams on resilience practices, documentation, and incident improvements.
Requirements
- 6–9 years of software engineering experience.
- Experience with distributed systems and microservices.
- Exposure to AWS, Azure, and GCP environments.
- Hands-on Kubernetes experience, including EKS, AKS, or GKE.
- Strong Java and Spring Boot skills.
- Understanding of distributed-system failures, resilience patterns, SLOs, and SLIs.
- Exposure to chaos-engineering tools such as Chaos Mesh, Gremlin, or Harness.
- Experience with Grafana, Prometheus, and LogDNA.
- Preferred qualifications include multi-region architecture, service mesh, performance testing, and Snyk experience.
- Good communication skills.
Benefits
- Flexible work arrangements supported by a distributed-workforce model.
- Equal opportunity workplace with reasonable accommodation for applicants with disabilities.
Tech Stack
Categories
Site Reliability
