15 hours ago
Gurgaon, IndiaStaff+
Responsibilities
- Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
- Design and implement scalable engineering solutions that reduce operational toil and embed reliability into delivery workflows.
- Shape engineering standards, reusable patterns, and frameworks across the SRE practice.
- Lead complex incident response, systemic remediation, and post-incident learning.
- Mentor and coach less senior engineers in reliability engineering, automation, observability, and SRE principles.
- Collaborate with engineering, platform, and operations teams to embed reliability and governance controls.
- Communicate engineering status, risks, and recommendations to senior stakeholders and leadership forums.
- Contribute metrics on service health, pipeline performance, automation coverage, and improvement progress to monthly operational reviews.
Requirements
- 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
- Strong hands-on experience across cloud, automation, observability, and CI/CD, including reliability solutions designed at scale.
- Deep experience operating cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
- Experience with Infrastructure-as-Code, reusable IaC patterns, landing zone components, and automated network controls.
- Strong knowledge of cloud networking, account management, identity primitives, policy enforcement, hybrid and cloud network architectures.
- Experience with incident response, SLO/SLI practices, telemetry pipelines, ingestion controls, observability cost management, and operational automation.
- Hands-on experience with identity platforms, secrets management, OIDC, workload identity, dynamic credentials, Zero Trust, and least-privilege adoption.
- Experience embedding security tooling in CI/CD pipelines and designing policy-as-code and secure-by-default controls.
- Strong scripting experience, such as Python, and hands-on experience with enterprise observability platforms such as Splunk or Datadog.
- Preferred qualifications include experience in federated, multi-cloud, or large enterprise environments; Docker and Kubernetes; professional cloud certification; OPA or Sentinel; AIOps; ServiceNow or PagerDuty; FinOps; and cloud policy-as-code tools.
- Strong technical communication, stakeholder engagement, mentoring, and cross-team leadership skills.
Benefits
- Hybrid or on-site work model, with regular office or client presence typically around 50% of working time.
- Participation in an on-call rotation and leadership of incident response.
- Occasional travel may be required for team or stakeholder engagement.
- In-person collaboration, mentorship, and professional development in a dynamic, collaborative environment.
Tech Stack
Categories
DevOpsSite Reliability
About BCG
Boston Consulting Group provides management consulting, technology, and design services to large enterprises, governments, and nonprofits, helping with strategy, digital transformation, and operations. The firm works on a project-based, fee-for-service model and also incubates products and ventures through BCG X. Founded in 1963 and headquartered in Boston, it is a global partnership with offices worldwide.
