
Lead Site Reliability Engineer - Imunify Reliability Platform (remote work)
CloudLinux18 days ago
Remote, EMEA +6 moreStaff+
Responsibilities
- Define and implement SLI, SLO, error-budget, ownership, and tiering frameworks for approximately 70 components.
- Design and build a push-based, sampled, privacy-constrained telemetry pipeline for fleet and service indicators.
- Extend agent-side and service-side instrumentation using Python, Go, and Rust.
- Consolidate dashboards, queries, and reporting into a maintainable set of observability instruments.
- Build symptom-based, SLO-anchored alerting with page, ticket, and dashboard tiers, owners, runbooks, and documented failure modes.
- Create component ownership maps, routing, severity matrices, acknowledgement SLAs, follow-the-sun rota practices, and handoff protocols.
- Coach product squads to carry their own pagers while operating the reliability platform and improving incident practices.
- Establish incident command and blameless postmortem processes with timelines, ownership, and action-item follow-through.
Requirements
- Substantial production-engineering or SRE experience, including defining an SLO framework rather than inheriting one.
- Strong Python skills and ability to read and modify Go or Rust.
- Practical experience with Prometheus/OpenMetrics, Grafana, an Alertmanager-class routing layer, and ClickHouse or an equivalent columnar store.
- Experience debugging distributed systems on bare metal and long-lived hosts, with limited reliance on Kubernetes.
- Production-scale configuration management and CI experience with Ansible, GitLab CI, Jenkins, or close equivalents.
- Ability to design push telemetry and sampling for machines that cannot be directly scraped, including handling clock skew, partial reporting, cardinality, and privacy constraints.
- Strong asynchronous written communication and ability to align engineering teams on health definitions.
- Security-product experience with WAF, EDR, antivirus, or vulnerability management is valuable.
- Experience with SOC 2 CC7.x, ISO 27001 A.8.16, or NIST SP 800-137 continuous monitoring is valuable.
- Experience with OpenTelemetry, eBPF, Sentry, cost- and cardinality-aware telemetry, Kubernetes, and agentic development tooling is valuable.
Benefits
- Fully remote work with flexible hours from any location worldwide.
- Professional-development opportunities, challenging projects, mentoring, and knowledge-exchange programs.
- 24 paid vacation days, 10 national holidays, and unlimited sick leave per year.
- Private medical insurance compensation.
- Co-working and gym/sports reimbursement.
- Opportunity to receive a reward for an innovative idea that the company can patent.
- Remote-first, async collaboration across nine time zones with weekly and monthly syncs and quarterly architecture summits.
Tech Stack
Categories
Site Reliability