BCG

Site Reliability Engineer Senior Manager

BCG
Apply
24 hours ago
Heredia, Costa RicaSenior / Staff+

Responsibilities

  • Run and continuously improve reliability engineering systems, automation, pipelines, observability, and operational tooling.
  • Design scalable automation and reliability solutions that reduce operational toil and embed reliability into delivery workflows.
  • Develop engineering standards, reusable frameworks, infrastructure-as-code patterns, landing-zone components, and governance controls across cloud providers.
  • Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
  • Design and operate telemetry pipelines, ingestion controls, observability signals, SLO/SLI practices, and observability cost-management processes.
  • Drive identity, secrets-management, Zero Trust, least-privilege, policy-as-code, secure engineering, network-control, and network-observability adoption.
  • Mentor and coach engineers and collaborate across engineering, platform, and operations teams.
  • Communicate engineering status, risks, and recommendations to senior stakeholders and leadership forums.

Requirements

  • 8–10 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
  • Strong hands-on experience across cloud, automation, observability, and CI/CD, including operation of cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
  • Deep knowledge of cloud networking, account management, identity primitives, policy enforcement, hybrid and cloud network architectures, and observability platforms.
  • Experience designing reusable Infrastructure-as-Code patterns, automated network controls, policy-as-code controls, and secure-by-default patterns.
  • Strong scripting experience, such as Python, and experience designing and operating telemetry pipelines, ingestion controls, signals, and ops automation.
  • Experience leading incident response, driving SLO/SLI practices, and implementing systemic reliability improvements across multiple teams.
  • Deep hands-on experience with enterprise observability platforms such as Splunk or Datadog, identity platforms such as Entra ID, and secrets-management tools such as HashiCorp Vault.
  • Experience designing OIDC, workload identity, dynamic credential, Zero Trust, least-privilege, and security-tooling patterns.
  • Preferred qualifications include experience in federated, multi-cloud, or large enterprise environments; Docker and Kubernetes; professional cloud certification; OPA or Sentinel; AIOps and event correlation; ServiceNow or PagerDuty; FinOps; cloud policy tools; and network reliability and security patterns.

Benefits

  • Hybrid or on-site work model.
  • Participation in an on-call rotation and leadership of incident response.
  • Occasional travel may be required for team or stakeholder engagement.
  • Equal Opportunity Employer with consideration for applicants protected under applicable law.

Categories

DevOpsSite Reliability
BCG

About BCG

10,000+ employees

Boston Consulting Group provides management consulting, technology, and design services to large enterprises, governments, and nonprofits, helping with strategy, digital transformation, and operations. The firm works on a project-based, fee-for-service model and also incubates products and ventures through BCG X. Founded in 1963 and headquartered in Boston, it is a global partnership with offices worldwide.

Contact me