3 months ago
Lisbon, PortugalSenior / Staff+
Responsibilities
- Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
- Design scalable engineering solutions that reduce operational toil and embed reliability into delivery workflows.
- Shape SRE engineering standards, reusable patterns, frameworks, and governance across teams.
- Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
- Mentor and coach less senior engineers across reliability engineering, automation, observability, and SRE practices.
- Drive cross-team collaboration across engineering, platform, and operations functions.
- Design and operate telemetry pipelines, ingestion controls, observability signals, SLO/SLI practices, and cost-management processes.
- Design cloud, identity, security, network, policy-as-code, and Infrastructure-as-Code controls across multiple providers.
- Communicate engineering status, risks, metrics, and recommendations to senior stakeholders and leadership forums.
Requirements
- 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
- Strong hands-on experience with cloud, automation, observability, CI/CD, and reliability solutions at scale.
- Deep experience operating infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
- Experience with Infrastructure-as-Code, reusable IaC patterns, landing zones, cloud networking, account management, identity, and policy enforcement.
- Strong scripting experience, such as Python, and experience designing and operating CI/CD pipelines.
- Experience leading incident response, driving systemic improvement, and implementing SLO/SLI practices across multiple teams.
- Deep hands-on experience with enterprise observability platforms, telemetry pipelines, ingestion controls, observability cost management, synthetic checks, and alerting.
- Deep hands-on experience with identity platforms, secrets management, OIDC, workload identity, dynamic credentials, Zero Trust, and least-privilege adoption.
- Experience embedding security tooling in CI/CD pipelines and designing policy-as-code and secure-by-default controls.
- Experience with hybrid and cloud network architectures, automated network controls, Zero Trust segmentation, and network observability.
- Preferred qualifications include federated or large-enterprise experience, Docker, Kubernetes, professional-level cloud certification, OPA, Sentinel, AIOps, event correlation, ServiceNow, PagerDuty, cloud FinOps, chargeback tooling, Alibaba Cloud architecture, and engineering communities of practice.
- Strong technical communication, stakeholder engagement, and ability to lead complex observability incidents and capacity reviews.
Benefits
- Hybrid or onsite work model.
- Participation in an on-call rotation and leadership of incident response.
- Occasional travel may be required for team or stakeholder engagement.
Tech Stack
Categories
Site Reliability
