
Senior DevOps Engineer
Qualified Health PBC3 months ago
Palo Alto, CA, USASenior
Base Salary
$170k - $220k/yr
Responsibilities
- Partner with engineering teams on production readiness, deployment patterns, failure modes, resource requirements, and rollback strategies.
- Design and maintain multi-cloud observability infrastructure covering metrics, logging, distributed tracing, dashboards, alerting policies, and SLIs/SLOs.
- Lead production incident response, root cause analysis, hotfix coordination, on-call rotations, runbooks, and postmortems.
- Provide operational support for deployments, production debugging, developer experience, and release processes.
- Design and maintain zero-trust network architectures and secure connectivity across multi-cloud environments and tenant boundaries.
- Build and improve CI/CD pipelines and release processes.
- Develop operational automation and tooling with Python and Terraform.
- Manage production Kubernetes workloads, including cluster troubleshooting, networking, resource optimization, and workload reliability.
- Operate Temporal workflows in production, including monitoring, scaling, and troubleshooting long-running executions.
- Collaborate with security and compliance teams on HIPAA and HITRUST controls.
Requirements
- 6+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, including at least 3 years directly managing production workloads.
- Strong Terraform proficiency, including module development, state management, and multi-environment architectures.
- Deep experience operating production Kubernetes environments and hands-on experience with Google Cloud Platform and Microsoft Azure.
- Strong networking and security knowledge, including zero-trust architectures, network segmentation, private connectivity, identity-based access controls, and secrets management.
- Production experience with Temporal or comparable workflow orchestration systems.
- Strong Python proficiency for automation, tooling, and operational scripting.
- Experience designing and operating observability stacks and leading incident response processes.
- Experience improving production readiness and release practices with engineering teams.
- Excellent written communication skills for runbooks, postmortems, and release documentation.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience.
- Preferred qualifications include healthcare and HIPAA experience, HITRUST familiarity, production LLM/agentic workflow/RAG experience, GitOps experience with Rancher Fleet, ArgoCD, or Flux, multi-tenant SaaS infrastructure experience, chaos engineering or reliability testing experience, and prior founding or early SRE/platform experience.
Benefits
- Competitive salary with equity packages.
- Medical, dental, and vision insurance.
- Flexible working hours.
- Hybrid work options.
- Inclusive environment focused on creativity and innovation.
Categories
DevOpsSite Reliability