BCG

Global IT Site Reliability Engineer Senior Manager

BCG
Apply
3 months ago
Lisbon, PortugalSenior / Staff+

Responsibilities

  • Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
  • Design scalable engineering solutions that reduce operational toil and embed reliability into delivery workflows.
  • Shape SRE engineering standards, reusable patterns, frameworks, and governance across teams.
  • Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
  • Mentor and coach less senior engineers across reliability engineering, automation, observability, and SRE practices.
  • Drive cross-team collaboration across engineering, platform, and operations functions.
  • Design and operate telemetry pipelines, ingestion controls, observability signals, SLO/SLI practices, and cost-management processes.
  • Design cloud, identity, security, network, policy-as-code, and Infrastructure-as-Code controls across multiple providers.
  • Communicate engineering status, risks, metrics, and recommendations to senior stakeholders and leadership forums.

Requirements

  • 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
  • Strong hands-on experience with cloud, automation, observability, CI/CD, and reliability solutions at scale.
  • Deep experience operating infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
  • Experience with Infrastructure-as-Code, reusable IaC patterns, landing zones, cloud networking, account management, identity, and policy enforcement.
  • Strong scripting experience, such as Python, and experience designing and operating CI/CD pipelines.
  • Experience leading incident response, driving systemic improvement, and implementing SLO/SLI practices across multiple teams.
  • Deep hands-on experience with enterprise observability platforms, telemetry pipelines, ingestion controls, observability cost management, synthetic checks, and alerting.
  • Deep hands-on experience with identity platforms, secrets management, OIDC, workload identity, dynamic credentials, Zero Trust, and least-privilege adoption.
  • Experience embedding security tooling in CI/CD pipelines and designing policy-as-code and secure-by-default controls.
  • Experience with hybrid and cloud network architectures, automated network controls, Zero Trust segmentation, and network observability.
  • Preferred qualifications include federated or large-enterprise experience, Docker, Kubernetes, professional-level cloud certification, OPA, Sentinel, AIOps, event correlation, ServiceNow, PagerDuty, cloud FinOps, chargeback tooling, Alibaba Cloud architecture, and engineering communities of practice.
  • Strong technical communication, stakeholder engagement, and ability to lead complex observability incidents and capacity reviews.

Benefits

  • Hybrid or onsite work model.
  • Participation in an on-call rotation and leadership of incident response.
  • Occasional travel may be required for team or stakeholder engagement.

Categories

Site Reliability
BCG

About BCG

10,000+ employees
Contact me