11 months ago
Bengaluru, IndiaSenior
Responsibilities
- Own uptime, reliability, performance, scalability, and cost efficiency for services running on AWS and Kubernetes/EKS.
- Design and implement self-healing infrastructure and AI-agent automation for operational workflows.
- Build LLM-powered tools for alert triage, incident summarization, root-cause analysis, and runbook automation.
- Manage Kubernetes deployments, autoscaling, resource optimization, and cluster reliability.
- Build and evolve metrics, dashboards, logging, and tracing observability systems.
- Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
- Automate infrastructure with Terraform and CI/CD pipelines.
- Lead incident response, postmortems, continuous reliability improvements, and chaos engineering practices.
Requirements
- At least 5 years of experience in SRE, DevOps, or platform engineering.
- Strong hands-on experience with AWS infrastructure at scale and production-grade Kubernetes clusters.
- Ability to debug complex distributed systems under pressure.
- Strong coding skills in Python or Go for building internal platforms and tools.
- Experience implementing monitoring, alerting, and incident management systems.
- Experience with LLM APIs such as the OpenAI API is a bonus.
- Familiarity with LangChain, AutoGen, AI agents for DevOps/SRE workflows, RAG systems, vector databases, AIOps, or intelligent automation is beneficial.
Benefits
- Growth and learning opportunities through challenging work with direct customer impact.
- Welcoming and positive work environment at a high-growth Platform as a Service company.
- Annual security training and adherence to Saviynt information security and privacy policies are required.
