IonQ

Senior Staff Service Reliability and Operational Intelligence Engineer

IonQ
Apply
4 hours ago
Santa Clara, CA, USAStaff+
H1B Sponsor

Base Salary

$188k - $270k/yr

Responsibilities

  • Own the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production.
  • Define and govern the New Service Introduction framework and production-readiness reviews covering architecture, security, resilience, capacity, observability, supportability, and releases.
  • Establish service ownership standards for catalogs, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness.
  • Lead the architecture and evolution of the shared observability platform, including standards for logs, metrics, distributed traces, profiles, dashboards, alerts, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls.
  • Own reliability governance involving SLIs, SLOs, error budgets, service health, customer impact, and escalation mechanisms.
  • Advance incident-management maturity through severity classification, incident command, communications, automated evidence collection, post-incident reviews, remediation tracking, and systemic fixes.
  • Lead capacity and efficiency management across demand forecasting, cloud and Kubernetes capacity, performance testing, scaling thresholds, headroom, rightsizing, and capacity-risk reviews.
  • Design and govern AIOps and secure AI-agent workflows for event correlation, alert-noise reduction, predictive detection, root-cause analysis, autonomous triage, remediation, and controlled self-healing.
  • Improve on-call effectiveness through rotation design, operational-readiness standards, escalation policies, diagnostic automation, alert-quality management, and follow-the-sun handoffs.
  • Provide hands-on technical leadership during major incidents, reliability investigations, architectural reviews, resilience exercises, and critical service launches.
  • Use operational, incident, service-level, capacity, change, and automation data to prioritize continuous improvement.

Requirements

  • 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Experience owning observability architecture and governing metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Experience establishing reliability and operational-readiness standards for business-critical services.
  • Hands-on experience with failure experiments, disaster-recovery exercises, and validated service failovers.
  • Experience commanding SEV1 or SEV2 incidents, coordinating technical and executive communications, and driving systemic remediation.
  • Experience delivering measurable reliability outcomes involving availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
  • Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
  • Strong software engineering and automation skills using Python or Go, infrastructure as code, and modern delivery toolchains.
  • Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond one team.
  • Ability to influence cross-functional stakeholders and deliver complex initiatives without direct management authority.
  • Preferred experience with risk prioritization, AIOps, autonomous remediation, AI-agent integrations, FinOps, load balancing, global traffic management, progressive delivery, and global follow-the-sun on-call models.
  • U.S. Person status, required authorization, or an applicable license exception may be necessary to access export-controlled technology and perform certain government-contract work.

Benefits

  • Comprehensive medical, dental, and vision plans.
  • Matching 401(k), unlimited PTO, paid holidays, parental/adoption leave, legal insurance, and a home technology stipend.
  • Located in Santa Clara, California, with up to 25% travel.
  • Equal opportunity workplace with accessibility and inclusion commitments.

Categories

DevOpsSite Reliability
IonQ

About IonQ

1,001-5,000 employees

IonQ, Inc. [NYSE: IONQ] is the world’s leading quantum platform and foundry - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the forthcoming IonQ Tempo, will be the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, AstraZeneca, and NVIDIA achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance. Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Toronto, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before. Learn more at IonQ.com.