BCG

Site Reliability Engineer Director

BCG
Apply
1 day ago
Heredia, Costa RicaStaff+

Responsibilities

  • Identify systemic platform and service risks, failure modes, and scalable engineering mitigations.
  • Embed operational activities into delivery models through automation, CI/CD integration, and event-driven workflows.
  • Design reliability patterns across cloud, network, identity, security, and observability domains.
  • Drive enterprise observability, AIOps, event correlation, anomaly detection, noise reduction, and automated remediation.
  • Lead secretless, least-privilege, secure-by-default, Zero Trust, and segmented network transformations.
  • Define secure engineering, cloud platform, identity, network, and automation strategies and standards.
  • Build reusable CI/CD, Terraform, governance, validation, and operational automation modules and frameworks.
  • Translate incident learnings and governance requirements into preventative controls and repeatable engineering standards.
  • Provide technical leadership, mentor engineers, and influence architecture and practices across multiple teams and business units.

Requirements

  • 10–12+ years of experience in Site Reliability Engineering, Platform Engineering, or related fields.
  • Strong hands-on experience with AWS, Azure, Terraform, Datadog or Splunk, and Python-based automation.
  • Deep understanding of distributed systems, failure modes, resilience patterns, telemetry pipelines, observability architecture, and operations automation.
  • Expertise in multi-cloud platform engineering, landing zones, governance, networking, Infrastructure as Code, and enterprise-scale cloud adoption across AWS, Azure, GCP, and/or Alibaba Cloud.
  • Expertise in enterprise identity, access, secrets engineering, OIDC, SAML, workload identity, and modern identity architecture.
  • Expertise in secure engineering architecture, policy-as-code, threat modeling, secure SDLC practices, and shift-left transformation.
  • Expertise in enterprise network architecture, hybrid and cloud networking, segmentation, observability, and security controls.
  • Proven experience driving SLO-driven engineering, signal-driven automation, Zero Trust, secretless transformation, governance, and adoption at scale.
  • Strong stakeholder engagement, technical communication, systems thinking, and technical leadership skills.
  • Preferred experience with Entra ID, HashiCorp Vault, compliance requirements such as PCI, HIPAA, and SOX, cloud FinOps, and large federated organizations.
  • Preferred experience setting observability, identity, secure engineering, network engineering, and automation governance strategies and standards.

Benefits

  • Hybrid or on-site work model.
  • Occasional travel may be required for team or stakeholder engagement.
  • Senior individual-contributor role with broad cross-organizational influence and a balance of hands-on technical leadership and strategic direction.

Tech Stack

Categories

DevOpsSite Reliability
BCG

About BCG

10,000+ employees

Boston Consulting Group provides management consulting, technology, and design services to large enterprises, governments, and nonprofits, helping with strategy, digital transformation, and operations. The firm works on a project-based, fee-for-service model and also incubates products and ventures through BCG X. Founded in 1963 and headquartered in Boston, it is a global partnership with offices worldwide.

Contact me