HostPapa Inc.

Site Reliability Engineer

HostPapa Inc.
Apply
23 days ago
Remote, WorldwideMid Level

Responsibilities

  • Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services.
  • Design and operate observability across metrics, logs, and traces using monitoring and logging tools.
  • Develop alerting strategies and dashboards for platform and business health.
  • Design and maintain high-availability, redundancy, failover, and disaster recovery architectures.
  • Conduct capacity planning, load testing, and performance optimization.
  • Lead production incident coordination, communication, service restoration, and blameless postmortems.
  • Improve Kubernetes platform reliability through health checks, autoscaling, rollout safety, and resilience testing.
  • Improve deployment safety, rollback strategies, automation, and operational processes with engineering and DevOps teams.
  • Maintain runbooks and operational documentation and promote SRE best practices.
  • Support additional projects and tasks needed by the team and business.

Requirements

  • At least 3 years of experience as an SRE, DevOps Engineer, or Production Engineer with strong production-system ownership.
  • Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms.
  • Hands-on experience with Datadog, Grafana, Elasticsearch, and/or Kibana.
  • Strong understanding of Linux, networking, and distributed-systems fundamentals.
  • Experience with containerized environments such as Docker and Kubernetes.
  • Strong scripting and automation skills using Python and/or Bash.
  • Experience participating in production on-call rotations and incident response.
  • Strong written and spoken English.
  • Experience defining SLIs/SLOs and managing error budgets at scale is a plus.
  • Experience with hyperscale or service-provider-grade platforms is advantageous.
  • Cloud experience, preferably Azure; AWS and/or GCP experience is also valued.
  • Experience with hybrid or on-premises integrations is beneficial.
  • Familiarity with chaos engineering and resilience testing is an asset.

Benefits

  • Remote opportunity open to applicants globally, with priority for candidates based in Spain.
  • Competitive salary and flexible work arrangements supporting work/life balance.
  • Career advancement and professional development opportunities.
  • Accommodation is available throughout the hiring process for people with disabilities.

Tech Stack

Categories

Site Reliability
Contact me