
Site Reliability Engineer
ZEISS Group1 hour ago
Bengaluru, IndiaSenior
Responsibilities
- Act as the first-line technical owner for applications running across multiple regional locations.
- Perform deep service-level diagnostics using observability tooling, including Grafana, Loki, and Mimir.
- Work with databases and messaging bus technologies to identify service-level issues.
- Support incident response, root cause analysis, and service restoration.
- Operate and support workloads running on Azure-based Kubernetes platforms.
- Understand Kubernetes pod lifecycle, networking, storage, scaling, and failure scenarios.
- Collaborate with platform teams on cluster-level issues and application teams on service-level issues.
- Report unresolved bugs and issues to product and platform teams after extended diagnostics.
Requirements
- Bachelor’s degree in computer science or engineering, or equivalent work experience.
- At least 5 years of experience in an SRE role or similar.
- Hands-on experience with Kubernetes and Azure.
- Experience with MongoDB, MSSQ, pSQL, and Kafka.
- Experience with Helm, Kustomize, Git, Azure DevOps, ArgoCD, and Terraform.
- AZ-104, CKA, or equivalent experience.
- Strong problem-solving and critical-thinking skills, with a proactive and independent working style.
Tech Stack
Categories
Site Reliability