
Site Reliability Engineer - SaaSOps
ValGenesis6 months ago
Hyderābād, IndiaMid Level
Responsibilities
- Define and embed SRE best practices across the SaaS platform.
- Establish and maintain SLAs, SLIs, SLOs, and error budgets.
- Design and improve high-availability and disaster recovery strategies.
- Automate manual processes, manage incident response, and optimize system performance.
- Ensure tenant isolation and consistent performance within a DB-per-tenant architecture.
- Improve system resiliency across Azure and on-premises environments.
- Lead incident response, root cause analysis, and blameless postmortems.
- Translate incidents into systemic fixes and maintain operational runbooks.
- Design and maintain observability across cloud and on-premises environments.
Requirements
- At least 3 years of hands-on Site Reliability Engineering experience supporting production-grade, cloud-native enterprise software platforms or applications.
- Prior experience as a DevOps engineer, cloud system administrator, or software developer.
- Strong proficiency with scripting languages such as Python and PowerShell.
- Deep hands-on production experience with Microsoft Azure.
- Experience with Terraform, Ansible, and Kubernetes internals, including networking, scheduling, scaling, and resource management.
- Production experience tuning and optimizing PostgreSQL performance.
- Hands-on experience with Azure Monitor, Application Insights, and Log Analytics.
- Experience implementing and managing Prometheus and Grafana for Kubernetes and on-premises monitoring.
- Ability to turn metrics, logs, and traces into actionable reliability and performance improvements.
- Experience troubleshooting and improving CI/CD pipelines.
- Understanding of GitOps principles for controlled and auditable deployments and infrastructure changes.
Benefits
- Chennai, Hyderabad, and Bangalore office roles are onsite five days per week.
Tech Stack
Categories
Site Reliability