3 months ago
Gurgaon, IndiaSenior / Staff+
Responsibilities
- Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
- Design and implement scalable solutions that eliminate operational toil and embed reliability into delivery workflows.
- Develop engineering standards, reusable patterns, and frameworks across the SRE practice.
- Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
- Mentor and coach less senior engineers across reliability engineering, automation, observability, and SRE principles.
- Drive cross-team collaboration across engineering, platform, and operations functions to embed reliability and governance controls.
- Communicate engineering status, risks, and recommendations to senior stakeholders and leadership forums.
- Design and operate telemetry pipelines, ingestion controls, observability cost-management practices, cloud infrastructure, identity controls, security tooling, and network architectures.
- Drive SLO/SLI practices, Zero Trust and least-privilege adoption, policy-as-code controls, and secure-by-default patterns across multiple teams.
Requirements
- 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
- Strong hands-on experience across cloud, automation, observability, and CI/CD, including experience designing reliability solutions at scale.
- Deep experience operating cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
- Experience with Infrastructure-as-Code, CI/CD pipelines, scripting, incident response, telemetry pipelines, ingestion controls, SLOs, SLIs, synthetic checks, alerts, and operations automation.
- Experience with cloud networking, account management, identity primitives, policy enforcement, hybrid and cloud network architectures, and automated network controls.
- Deep hands-on experience with enterprise observability platforms, identity platforms, secrets management, and security tooling embedded in CI/CD pipelines.
- Experience designing OIDC, workload identity, dynamic credential, policy-as-code, Zero Trust, least-privilege, and secure-by-default patterns.
- Experience driving cloud platform engineering standards, governance, secure engineering adoption, network observability, and SLO/SLI practices across multiple teams.
- Strong scripting experience, such as Python, and strong stakeholder engagement and technical communication skills.
- Preferred qualifications include large-enterprise or federated multi-cloud experience; Docker and Kubernetes familiarity; professional-level cloud certification; experience with OPA, Sentinel, AIOps, event correlation, ServiceNow, PagerDuty, FinOps, chargeback tooling, Alibaba Cloud architecture, AWS Service Control Policies, and Azure Policy.
Benefits
- Hybrid or on-site work model.
- Participation in an on-call rotation and leadership of incident response are expected.
- Occasional travel may be required for team or stakeholder engagement.
Tech Stack
Categories
Site Reliability
