1 month ago
Remote, SpainSenior
Responsibilities
- Design, implement, and maintain highly available, scalable, and resilient systems.
- Build automation, self-service tooling, and software to reduce operational toil and improve reliability.
- Own observability practices including monitoring, alerting, logging, tracing, dashboards, and synthetic testing.
- Lead incident response, participate in on-call rotations, and drive blameless post-mortems.
- Define, implement, and track SLIs, SLOs, and error budgets.
- Use Terraform, Flux, and GitHub Actions for infrastructure as code, GitOps, and CI/CD automation.
- Provide reliability expertise during system design reviews and influence architectural decisions.
- Document processes, create runbooks, mentor engineers, and develop AI-assisted operational workflows with appropriate human oversight.
Requirements
- Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role.
- Ability to build mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability.
- Strong troubleshooting, incident response, and root cause analysis skills.
- Experience defining and using SLIs, SLOs, and error budgets.
- Experience with Kubernetes platforms including Amazon EKS, service meshes such as Istio, and AWS infrastructure and services.
- Experience with identity and access management systems including Auth0 and AWS IAM.
- Experience with GitOps workflows and infrastructure automation using Flux and Terraform.
- Experience with observability platforms and practices, CI/CD systems, application logging, and distributed-system debugging.
- Ability to build and maintain automation and tooling using one or more of Python, Go, or Bash.
- Ability to lead complex incident follow-up, communicate clearly, and apply systems thinking to reliability improvements.
- Ability to use and validate AI-assisted troubleshooting, root cause analysis, documentation, and operational workflows.
- Experience with feature-flagging platforms such as LaunchDarkly is a bonus.
Benefits
- Fully remote work arrangement.
- Annual compensation range of 70,000–115,000 EUR.
- Opportunity to shape SRE practices and improve operational excellence within a cloud-based enterprise software company.
Tech Stack
Categories
Site Reliability
