
Senior Site Reliability Engineer
ZEISS Group2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Define and measure service reliability using SLIs, SLOs, and service-level objectives aligned with business and user expectations.
- Enable development teams to release software and features quickly while maintaining agreed operational performance and risk levels.
- Design and maintain observability, monitoring, logging, and alerting for applications and managed cloud components.
- Build playbooks and runbooks for troubleshooting, incident response, and issue investigation.
- Set up and improve automation and deployment pipelines for continuously releasing software applications.
- Monitor system performance and reliability, respond to incidents, conduct post-incident reviews, and improve system resilience.
- Create and maintain documentation for systems, processes, disaster recovery plans, and operating procedures.
- Advocate for SRE principles and best practices while collaborating with developers, product owners, and cloud platform engineers.
Requirements
- In-depth knowledge of system architecture, networking, and distributed systems.
- Expertise designing and implementing reliable, scalable, and fault-tolerant systems.
- Proficiency with monitoring, alerting, and logging systems for Kubernetes and similar container orchestrators.
- Hands-on experience with incident response, troubleshooting, and post-mortem analysis.
- Proficiency in infrastructure automation and monitoring scripting or coding, including tools such as Terraform.
- Knowledge of deployment processes and strategies and cloud-based disaster recovery planning and execution.
- Strong communication and cross-functional collaboration skills.
- Willingness to stay current with site reliability engineering tools, technologies, and practices.
Tech Stack
Categories
Site Reliability