IonQ

Staff Service Reliability and Operational Intelligence Engineer

IonQ
Apply
10 hours ago
Santa Clara, CA, USAStaff+
H1B sponsor

Responsibilities

  • Define the technical strategy and roadmap for operational excellence, production readiness, and reliability across services and environments.
  • Establish service ownership, new-service-introduction, operational-readiness, and reliability governance standards.
  • Lead the architecture and evolution of observability platforms covering logs, metrics, traces, profiles, dashboards, alerting, synthetic monitoring, and telemetry governance.
  • Own SLI, SLO, error-budget, service-health, incident-management, post-incident review, and systemic-remediation practices.
  • Lead capacity forecasting, performance testing, scaling, resource optimization, cloud and Kubernetes capacity management, and capacity-risk reviews.
  • Design AIOps and secure AI-agent workflows for event correlation, anomaly detection, root-cause analysis, triage, remediation, and self-healing.
  • Provide hands-on technical leadership during major incidents, reliability investigations, resilience exercises, architecture reviews, and critical launches.
  • Improve global on-call effectiveness, escalation practices, diagnostic automation, alert quality, and follow-the-sun handoffs.
  • Use operational and incident data to prioritize continuous reliability and efficiency improvements.

Requirements

  • 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Experience owning observability architecture and governance for metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Experience establishing operational-readiness standards, conducting failure experiments and disaster-recovery exercises, and validating service failovers.
  • Experience commanding SEV1 or SEV2 incidents and driving root causes through systemic remediation.
  • Experience delivering measurable reliability outcomes involving availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
  • Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
  • Strong software engineering and automation skills using Python or Go, infrastructure as code, and modern delivery toolchains.
  • Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and broadly adopted mechanisms.
  • Ability to influence cross-functional stakeholders and deliver complex initiatives without direct management authority.
  • Preferred qualifications include experience with AIOps, autonomous remediation, Amazon Bedrock AgentCore or comparable agentic frameworks, operational-platform integrations, FinOps, high-availability traffic management, progressive delivery, and global on-call models.

Benefits

  • The package includes base compensation, bonus, equity, comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO, paid holidays, parental and adoption leave, legal insurance, and a home technology stipend.
  • The role is based in the Santa Clara, California office with the option to work remotely a few days per week.
  • Travel is expected up to 25%.

Categories

Site Reliability
IonQ

About IonQ

1,001-5,000 employees

IonQ builds trapped-ion quantum computers and offers access to them via major cloud platforms and direct systems for enterprise and research users. The public company (NYSE: IONQ), founded in 2015 and headquartered in College Park, Maryland, sells quantum hardware, cloud services, and development tools used in areas like materials modeling, optimization, and machine learning. Its systems are available through AWS Braket, Microsoft Azure Quantum, and Google Cloud Marketplace.

Contact me