Saviynt

Senior Site Reliability Engineer

Saviynt
Apply
11 months ago
Bengaluru, IndiaSenior

Responsibilities

  • Own uptime, reliability, performance, scalability, and cost efficiency for services running on AWS and Kubernetes/EKS.
  • Design and implement self-healing infrastructure and AI-agent automation for operational workflows.
  • Build LLM-powered tools for alert triage, incident summarization, root-cause analysis, and runbook automation.
  • Manage Kubernetes deployments, autoscaling, resource optimization, and cluster reliability.
  • Build and evolve metrics, dashboards, logging, and tracing observability systems.
  • Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
  • Automate infrastructure with Terraform and CI/CD pipelines.
  • Lead incident response, postmortems, continuous reliability improvements, and chaos engineering practices.

Requirements

  • At least 5 years of experience in SRE, DevOps, or platform engineering.
  • Strong hands-on experience with AWS infrastructure at scale and production-grade Kubernetes clusters.
  • Ability to debug complex distributed systems under pressure.
  • Strong coding skills in Python or Go for building internal platforms and tools.
  • Experience implementing monitoring, alerting, and incident management systems.
  • Experience with LLM APIs such as the OpenAI API is a bonus.
  • Familiarity with LangChain, AutoGen, AI agents for DevOps/SRE workflows, RAG systems, vector databases, AIOps, or intelligent automation is beneficial.

Benefits

  • Growth and learning opportunities through challenging work with direct customer impact.
  • Welcoming and positive work environment at a high-growth Platform as a Service company.
  • Annual security training and adherence to Saviynt information security and privacy policies are required.

Tech Stack

Categories

DevOpsSite Reliability
Saviynt

About Saviynt

1,001-5,000 employees
Contact me