20 days ago
Base Salary
$217k - $304k/yr
Responsibilities
- Lead reliability, scalability, performance, and operational excellence initiatives for critical user-facing systems.
- Design and guide highly available architectures covering failover, redundancy, graceful degradation, traffic management, and capacity planning.
- Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
- Build automation and tooling for safer deployments, incident response, remediation, and reliability guardrails.
- Lead complex incident response, blameless postmortems, root-cause analysis, and sustainable remediation.
- Define and promote standards for SLIs, SLOs, capacity management, release engineering, and operational maturity.
- Mentor SRE and software engineering teams and influence reliability culture across the organization.
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Experience supporting high-traffic, user-facing production environments and designing highly available systems.
- Deep understanding of distributed systems, networking, Linux systems, or cloud-native architectures.
- Strong programming skills in Go, Python, or similar languages.
- Strong understanding of observability systems, including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
- Experience with internet-scale systems, Kubernetes, containers, cloud infrastructure, modern deployment platforms, CDN optimization, edge reliability, traffic engineering, or global infrastructure is preferred.
- Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies is preferred.
- Experience leading large-scale incident response and operational transformation initiatives is preferred.
Benefits
- Global benefits supporting workspace needs, professional development, and caregiving.
- Family planning support, gender-affirming care, and mental health and coaching benefits.
- Private medical, dental, and vision benefits.
- Personal retirement savings account with matching contribution.
- Cycle to Work and Tax Saver schemes.
- Flexible vacation and paid volunteer time off.
- Generous paid parental leave.
- U.S.-based roles may include medical, dental, vision, and 401(k) benefits; interviews may be recorded, transcribed, and summarized by AI with an opt-out option.
Tech Stack
Categories
Site Reliability
About Reddit
Reddit builds a social news and discussion platform organized into user-created forums (“subreddits”) for consumers, creators, and communities. It monetizes primarily through advertising tools for brands and self-serve advertisers, plus a Premium subscription; developers can access data via its API. Founded in 2005 and headquartered in San Francisco, Reddit became a public company on the NYSE in 2024 and hosts over 100,000 active communities.
