10 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Own uptime, reliability, and performance for services running on AWS and Kubernetes EKS.
- Design and implement self-healing infrastructure using automation and AI agents.
- Build LLM-powered operational tooling for alert triage, incident summarization, root cause analysis, and runbook automation.
- Manage and scale Kubernetes workloads, including deployments, autoscaling, resource optimization, cluster reliability, and cost efficiency.
- Build and evolve observability systems covering metrics, dashboards, logs, and distributed tracing.
- Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
- Automate infrastructure with Terraform and CI/CD pipelines.
- Lead incident response, postmortems, continuous reliability improvements, and chaos engineering practices.
Requirements
- 8+ years of experience in SRE, DevOps, or Platform Engineering.
- Strong hands-on experience with AWS infrastructure at scale and production-grade Kubernetes clusters.
- Ability to debug complex distributed systems under pressure.
- Strong coding skills in Python or Go, including building internal platforms and tools.
- Experience implementing monitoring, alerting, and incident management systems.
- Preferred experience with LLM APIs such as the OpenAI API and agent frameworks such as LangChain and AutoGen.
- Preferred experience building AI agents for DevOps or SRE workflows, RAG systems, vector databases, or AIOps and intelligent automation systems.
Benefits
- Growth and learning opportunities in a high-growth Platform as a Service company.
- Work in a welcoming and positive environment on challenging work that directly impacts customers.
- Employees must follow information security and privacy policies and complete annual security training.
- Saviynt is an equal opportunity employer.
