ZEISS Group

Senior Site Reliability Engineer

ZEISS Group
Apply
2 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Define and measure service reliability using SLIs, SLOs, and service-level objectives aligned with business and user expectations.
  • Enable development teams to release software and features quickly while maintaining agreed operational performance and risk levels.
  • Design and maintain observability, monitoring, logging, and alerting for applications and managed cloud components.
  • Build playbooks and runbooks for troubleshooting, incident response, and issue investigation.
  • Set up and improve automation and deployment pipelines for continuously releasing software applications.
  • Monitor system performance and reliability, respond to incidents, conduct post-incident reviews, and improve system resilience.
  • Create and maintain documentation for systems, processes, disaster recovery plans, and operating procedures.
  • Advocate for SRE principles and best practices while collaborating with developers, product owners, and cloud platform engineers.

Requirements

  • In-depth knowledge of system architecture, networking, and distributed systems.
  • Expertise designing and implementing reliable, scalable, and fault-tolerant systems.
  • Proficiency with monitoring, alerting, and logging systems for Kubernetes and similar container orchestrators.
  • Hands-on experience with incident response, troubleshooting, and post-mortem analysis.
  • Proficiency in infrastructure automation and monitoring scripting or coding, including tools such as Terraform.
  • Knowledge of deployment processes and strategies and cloud-based disaster recovery planning and execution.
  • Strong communication and cross-functional collaboration skills.
  • Willingness to stay current with site reliability engineering tools, technologies, and practices.

Tech Stack

GrafanaKubernetesPrometheusTerraform

Categories

Site Reliability
ZEISS Group

About ZEISS Group

10,000+ employees
Contact me