24 hours ago
Heredia, Costa RicaSenior / Staff+
Responsibilities
- Run and continuously improve reliability engineering systems, automation, pipelines, observability, and operational tooling.
- Design scalable automation and reliability solutions that reduce operational toil and embed reliability into delivery workflows.
- Develop engineering standards, reusable frameworks, infrastructure-as-code patterns, landing-zone components, and governance controls across cloud providers.
- Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
- Design and operate telemetry pipelines, ingestion controls, observability signals, SLO/SLI practices, and observability cost-management processes.
- Drive identity, secrets-management, Zero Trust, least-privilege, policy-as-code, secure engineering, network-control, and network-observability adoption.
- Mentor and coach engineers and collaborate across engineering, platform, and operations teams.
- Communicate engineering status, risks, and recommendations to senior stakeholders and leadership forums.
Requirements
- 8–10 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
- Strong hands-on experience across cloud, automation, observability, and CI/CD, including operation of cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
- Deep knowledge of cloud networking, account management, identity primitives, policy enforcement, hybrid and cloud network architectures, and observability platforms.
- Experience designing reusable Infrastructure-as-Code patterns, automated network controls, policy-as-code controls, and secure-by-default patterns.
- Strong scripting experience, such as Python, and experience designing and operating telemetry pipelines, ingestion controls, signals, and ops automation.
- Experience leading incident response, driving SLO/SLI practices, and implementing systemic reliability improvements across multiple teams.
- Deep hands-on experience with enterprise observability platforms such as Splunk or Datadog, identity platforms such as Entra ID, and secrets-management tools such as HashiCorp Vault.
- Experience designing OIDC, workload identity, dynamic credential, Zero Trust, least-privilege, and security-tooling patterns.
- Preferred qualifications include experience in federated, multi-cloud, or large enterprise environments; Docker and Kubernetes; professional cloud certification; OPA or Sentinel; AIOps and event correlation; ServiceNow or PagerDuty; FinOps; cloud policy tools; and network reliability and security patterns.
Benefits
- Hybrid or on-site work model.
- Participation in an on-call rotation and leadership of incident response.
- Occasional travel may be required for team or stakeholder engagement.
- Equal Opportunity Employer with consideration for applicants protected under applicable law.
Tech Stack
Categories
DevOpsSite Reliability
About BCG
Boston Consulting Group provides management consulting, technology, and design services to large enterprises, governments, and nonprofits, helping with strategy, digital transformation, and operations. The firm works on a project-based, fee-for-service model and also incubates products and ventures through BCG X. Founded in 1963 and headquartered in Boston, it is a global partnership with offices worldwide.
