IonQ

Staff Site Reliability Engineer

IonQ
Apply
10 hours ago
Santa Clara, CA, USAStaff+
H1B sponsor

Base Salary

$152k - $228k/yr

Responsibilities

  • Own production reliability outcomes, service-level objectives, error budgets, and reliability standards for Tier-1 and Tier-2 services.
  • Design and operate observability platforms covering metrics, logs, distributed tracing, and profiles, while defining platform-wide instrumentation standards.
  • Govern SLO reviews, error-budget corrective actions, architecture decisions, and scaling decisions with service owners.
  • Design and execute chaos engineering experiments, disaster-recovery exercises, backup validation, failover testing, and recovery verification.
  • Lead the highest-severity incident response as incident commander, including serious SEV1/SEV2 and covered security incidents, and drive systemic fixes through post-incident reviews.
  • Establish on-call rotations, escalation paths, operational readiness practices, and follow-the-sun handoffs across regions.
  • Co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps.
  • Own reliability, capacity planning, rightsizing, and efficiency improvements for stateful and streaming services.
  • Build AI Ops workflows for autonomous triage, predictive alerting, remediation, and self-healing.
  • Mentor engineers, establish standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development around a reliability roadmap.

Requirements

  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Recent hands-on experience operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Experience instrumenting production systems and governing service-level objectives and error budgets.
  • Experience designing and executing failure experiments or disaster-recovery exercises with real failover validation.
  • Personal experience commanding serious SEV1/SEV2 incidents and driving root causes through systemic fixes.
  • Demonstrated measurable ownership of reliability outcomes such as availability, mean time to recovery, and error-budget adherence.
  • Evidence of multi-team technical leadership through standards, reviews, coaching, and mechanisms adopted beyond one service or team.
  • Preferred experience includes cloud security posture management, runtime vulnerability detection, workload protection, autonomous remediation, AIOps, capacity management, FinOps-based cost optimization, load balancing, global traffic management, and AI traffic management through an LLM gateway.
  • Preferred experience includes Amazon Bedrock Agent Core or equivalent agentic automation frameworks and the ability to integrate networking, security, and reliability into platform design decisions.

Benefits

  • The role is based in the Santa Clara, California office with the option to work remotely a few days per week.
  • Travel is expected to be up to 25%.
  • The compensation package includes base salary, bonus, equity, medical, dental, vision, matching 401(k), unlimited PTO, paid holidays, parental/adoption leave, legal insurance, and a home technology stipend.

Categories

Site Reliability
IonQ

About IonQ

1,001-5,000 employees

IonQ builds trapped-ion quantum computers and offers access to them via major cloud platforms and direct systems for enterprise and research users. The public company (NYSE: IONQ), founded in 2015 and headquartered in College Park, Maryland, sells quantum hardware, cloud services, and development tools used in areas like materials modeling, optimization, and machine learning. Its systems are available through AWS Braket, Microsoft Azure Quantum, and Google Cloud Marketplace.

Contact me