3 months ago
Heredia, Costa RicaSenior
Responsibilities
- Run and continuously improve reliability engineering systems, including automation, pipelines, observability, and operational tooling.
- Design scalable engineering solutions that reduce operational toil and embed reliability into delivery workflows.
- Create engineering standards, reusable patterns, frameworks, infrastructure-as-code components, and governance controls.
- Lead complex incident response, systemic remediation, post-incident learning, and operational reviews.
- Mentor and coach less senior engineers across reliability engineering, automation, observability, and SRE practices.
- Drive cross-team collaboration across engineering, platform, and operations functions.
- Communicate engineering status, risks, recommendations, service health metrics, pipeline performance, and improvement progress to stakeholders and leadership.
- Design and operate telemetry pipelines, ingestion controls, observability cost management, identity patterns, secure-by-default controls, and automated network controls.
Requirements
- 5–8 years of experience in Site Reliability Engineering, Platform Engineering, or related operational engineering disciplines.
- Strong hands-on experience across cloud, automation, observability, CI/CD, and reliability engineering at scale.
- Deep experience operating cloud infrastructure across at least two of AWS, Azure, GCP, or Alibaba Cloud.
- Experience with Infrastructure-as-Code, reusable cloud patterns, landing zones, cloud networking, account management, identity, and policy enforcement.
- Strong scripting experience, such as Python, and experience designing and operating automation and reliability solutions.
- Experience leading incident response, systemic improvement, SLO/SLI practices, observability signals, synthetic checks, alerts, and operations automation.
- Deep hands-on experience with enterprise observability platforms such as Splunk or Datadog and with telemetry and ingestion pipelines.
- Deep hands-on experience with identity platforms, secrets management, OIDC, workload identity, dynamic credentials, Zero Trust, and least-privilege adoption.
- Experience embedding security tooling in CI/CD pipelines, designing policy-as-code controls, and driving secure engineering adoption.
- Deep experience with hybrid and cloud network architectures, automated network controls, Zero Trust segmentation, and network observability.
- Preferred qualifications include federated, multi-cloud, or large-enterprise experience; Docker and Kubernetes familiarity; professional cloud certification; OPA or Sentinel; AIOps; ServiceNow or PagerDuty; FinOps; Alibaba Cloud architecture; and cloud policy tools such as AWS Service Control Policies or Azure Policy.
- Strong stakeholder engagement, technical communication, mentoring, engineering leadership, and understanding of identity, application security, network reliability, and observability risks.
Benefits
- Hybrid or on-site work model.
- Senior individual-contributor role with mentorship and cross-team influence.
- Participation in an on-call rotation and leadership of incident response.
- Occasional travel may be required for team or stakeholder engagement.
- Equal Opportunity Employer.
Tech Stack
Categories
Site Reliability
