Site Reliability Engineer
HostPapa Inc.23 days ago
Remote, WorldwideMid Level
Responsibilities
- Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services.
- Design and operate observability across metrics, logs, and traces using monitoring and logging tools.
- Develop alerting strategies and dashboards for platform and business health.
- Design and maintain high-availability, redundancy, failover, and disaster recovery architectures.
- Conduct capacity planning, load testing, and performance optimization.
- Lead production incident coordination, communication, service restoration, and blameless postmortems.
- Improve Kubernetes platform reliability through health checks, autoscaling, rollout safety, and resilience testing.
- Improve deployment safety, rollback strategies, automation, and operational processes with engineering and DevOps teams.
- Maintain runbooks and operational documentation and promote SRE best practices.
- Support additional projects and tasks needed by the team and business.
Requirements
- At least 3 years of experience as an SRE, DevOps Engineer, or Production Engineer with strong production-system ownership.
- Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms.
- Hands-on experience with Datadog, Grafana, Elasticsearch, and/or Kibana.
- Strong understanding of Linux, networking, and distributed-systems fundamentals.
- Experience with containerized environments such as Docker and Kubernetes.
- Strong scripting and automation skills using Python and/or Bash.
- Experience participating in production on-call rotations and incident response.
- Strong written and spoken English.
- Experience defining SLIs/SLOs and managing error budgets at scale is a plus.
- Experience with hyperscale or service-provider-grade platforms is advantageous.
- Cloud experience, preferably Azure; AWS and/or GCP experience is also valued.
- Experience with hybrid or on-premises integrations is beneficial.
- Familiarity with chaos engineering and resilience testing is an asset.
Benefits
- Remote opportunity open to applicants globally, with priority for candidates based in Spain.
- Competitive salary and flexible work arrangements supporting work/life balance.
- Career advancement and professional development opportunities.
- Accommodation is available throughout the hiring process for people with disabilities.
Tech Stack
Categories
Site Reliability