
Staff Site Reliability Engineer
Pivotal Health16 days ago
Remote, United States or Santa Monica, CA, USAStaff+
Base Salary
$230k - $250k/yr
Responsibilities
- Define the technical vision and roadmap for reliability, availability, scalability, and operational readiness, including service-level objectives.
- Design and implement resilient cloud architecture, deployment systems, networking, compute, storage, and shared infrastructure.
- Build cohesive observability through metrics, logs, traces, dashboards, and alerting.
- Establish incident management, on-call, postmortem, and disaster recovery practices and lead complex incident response.
- Create automation and internal tooling for deployment workflows, capacity management, infrastructure provisioning, production diagnostics, and other operational toil.
- Partner with software, data, and AI engineers to improve system design, production readiness, and failure handling.
- Support infrastructure security, access controls, audit trails, and compliance requirements for sensitive healthcare and financial data.
- Provide technical leadership through architecture reviews, mentoring, and guidance on infrastructure and operational engineering practices.
Requirements
- 8+ years of experience in site reliability engineering, infrastructure engineering, platform engineering, or operating large-scale production systems.
- Deep knowledge of cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, and modern deployment practices.
- Experience designing, building, and operating highly available systems in a fast-growing production environment.
- Strong understanding of observability, service-level objectives, capacity planning, incident management, disaster recovery, and performance engineering.
- Hands-on ability to debug complex issues across application, infrastructure, network, and data layers.
- Experience creating automation and internal tooling that reduces operational toil and improves developer productivity.
- Ability to influence architecture and engineering practices across teams without formal authority.
- Strong communication skills and pragmatic judgment when balancing reliability, delivery speed, complexity, and cost.
- Extra credit for experience with healthcare, financial, or other regulated data; HIPAA, PHI, SOC 2, data privacy, auditability, and security controls; production data-intensive, AI, or machine learning systems; multi-region architecture; high-volume asynchronous workflows; complex distributed processing; or introducing SRE practices at an early-stage or high-growth company.
Benefits
- Competitive compensation including equity, full health, dental, and vision coverage, a 401(k) retirement savings plan, flexible time off, and company-wide connection and events.
- Primarily hires around Los Angeles and New York hubs, with remote and hybrid flexibility varying by role and team.
- Employment requires authorization to work in the United States without current or future employer sponsorship.
Categories
Site Reliability