BCG

Senior Site Reliability Engineer

BCG
Apply
3 months ago
Heredia, Costa RicaSenior

Responsibilities

  • Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
  • Design scalable engineering solutions that reduce operational toil and embed reliability into delivery workflows.
  • Create engineering standards, reusable patterns, frameworks, infrastructure-as-code components, and governance controls.
  • Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
  • Mentor and coach less senior engineers across reliability engineering, automation, observability, and SRE practices.
  • Drive cross-team collaboration across engineering, platform, and operations functions.
  • Communicate engineering status, risks, recommendations, service health metrics, pipeline performance, and improvement progress to stakeholders and leadership.
  • Design and operate telemetry pipelines, ingestion controls, observability cost management, identity patterns, secure-by-default controls, and automated network controls.

Requirements

  • 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
  • Strong hands-on experience across cloud, automation, observability, CI/CD, and reliability engineering at scale.
  • Deep experience operating cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
  • Experience with Infrastructure-as-Code, reusable cloud patterns, landing zones, cloud networking, account management, identity, and policy enforcement.
  • Strong scripting experience, such as Python, and experience designing and operating automation and reliability solutions.
  • Experience leading incident response, systemic improvement, SLO/SLI practices, observability signals, synthetic checks, alerts, and operations automation.
  • Deep hands-on experience with enterprise observability platforms such as Splunk or Datadog and with telemetry and ingestion pipelines.
  • Deep hands-on experience with identity platforms, secrets management, OIDC, workload identity, dynamic credentials, Zero Trust, and least-privilege adoption.
  • Experience embedding security tooling in CI/CD pipelines, designing policy-as-code controls, and driving secure engineering adoption.
  • Deep experience with hybrid and cloud network architectures, automated network controls, Zero Trust segmentation, and network observability.
  • Preferred qualifications include federated, multi-cloud, or large-enterprise experience; Docker and Kubernetes familiarity; professional cloud certification; OPA or Sentinel; AIOps; ServiceNow or PagerDuty; FinOps; Alibaba Cloud architecture; and cloud policy tools such as AWS Service Control Policies or Azure Policy.
  • Strong stakeholder engagement, technical communication, mentoring, engineering leadership, and understanding of identity, application security, network reliability, and observability risks.

Benefits

  • Hybrid or on-site work model.
  • Senior individual-contributor role with mentorship and cross-team influence.
  • Participation in an on-call rotation and leadership of incident response.
  • Occasional travel may be required for team or stakeholder engagement.
  • Equal Opportunity Employer.

Categories

Site Reliability
BCG

About BCG

10,000+ employees
Contact me