CloudLinux

Lead Site Reliability Engineer - Imunify Reliability Platform (remote work)

CloudLinux
Apply
18 days ago
Remote, EMEA +6 moreStaff+

Responsibilities

  • Define and implement SLI, SLO, error-budget, ownership, and tiering frameworks for approximately 70 components.
  • Design and build a push-based, sampled, privacy-constrained telemetry pipeline for fleet and service indicators.
  • Extend agent-side and service-side instrumentation using Python, Go, and Rust.
  • Consolidate dashboards, queries, and reporting into a maintainable set of observability instruments.
  • Build symptom-based, SLO-anchored alerting with page, ticket, and dashboard tiers, owners, runbooks, and documented failure modes.
  • Create component ownership maps, routing, severity matrices, acknowledgement SLAs, follow-the-sun rota practices, and handoff protocols.
  • Coach product squads to carry their own pagers while operating the reliability platform and improving incident practices.
  • Establish incident command and blameless postmortem processes with timelines, ownership, and action-item follow-through.

Requirements

  • Substantial production-engineering or SRE experience, including defining an SLO framework rather than inheriting one.
  • Strong Python skills and ability to read and modify Go or Rust.
  • Practical experience with Prometheus/OpenMetrics, Grafana, an Alertmanager-class routing layer, and ClickHouse or an equivalent columnar store.
  • Experience debugging distributed systems on bare metal and long-lived hosts, with limited reliance on Kubernetes.
  • Production-scale configuration management and CI experience with Ansible, GitLab CI, Jenkins, or close equivalents.
  • Ability to design push telemetry and sampling for machines that cannot be directly scraped, including handling clock skew, partial reporting, cardinality, and privacy constraints.
  • Strong asynchronous written communication and ability to align engineering teams on health definitions.
  • Security-product experience with WAF, EDR, antivirus, or vulnerability management is valuable.
  • Experience with SOC 2 CC7.x, ISO 27001 A.8.16, or NIST SP 800-137 continuous monitoring is valuable.
  • Experience with OpenTelemetry, eBPF, Sentry, cost- and cardinality-aware telemetry, Kubernetes, and agentic development tooling is valuable.

Benefits

  • Fully remote work with flexible hours from any location worldwide.
  • Professional-development opportunities, challenging projects, mentoring, and knowledge-exchange programs.
  • 24 paid vacation days, 10 national holidays, and unlimited sick leave per year.
  • Private medical insurance compensation.
  • Co-working and gym/sports reimbursement.
  • Opportunity to receive a reward for an innovative idea that the company can patent.
  • Remote-first, async collaboration across nine time zones with weekly and monthly syncs and quarterly architecture summits.

Tech Stack

AnsibleClickHouseGitLab CI/CDGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRust

Categories

Site Reliability
CloudLinux

About CloudLinux

201-500 employees
Contact me