IonQ

Staff Service Reliability and Operational Intelligence Engineer

IonQ
Apply
11 hours ago
Santa Clara, CA, USAStaff+
H1B sponsor

Base Salary

$152k - $228k/yr

Responsibilities

  • Define the technical strategy, roadmap, and production-readiness standards for operational excellence across development, pre-production, and production environments.
  • Establish service ownership, service catalog, dependency mapping, runbook, support, escalation, recovery, and on-call standards.
  • Lead the architecture and governance of shared observability platforms covering logs, metrics, distributed traces, profiles, dashboards, alerts, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls.
  • Own reliability governance for SLIs, SLOs, error budgets, service health, customer impact, and escalation mechanisms.
  • Lead high-severity incident command, communications, evidence collection, blameless reviews, systemic remediation, and continuous improvement.
  • Manage capacity forecasting, performance testing, scaling, cloud and Kubernetes resources, rightsizing, headroom, and capacity risk.
  • Design AIOps and secure AI-agent workflows for anomaly detection, event correlation, predictive detection, root-cause analysis, triage, remediation, and self-healing.
  • Improve on-call operations through rotation design, escalation policies, diagnostic automation, alert-quality management, and follow-the-sun handoffs.
  • Provide hands-on technical leadership during incidents, reliability investigations, architecture reviews, resilience exercises, and critical service launches.

Requirements

  • 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep knowledge of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Experience owning observability architecture and governing metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Experience establishing reliability and operational-readiness standards for business-critical services.
  • Hands-on experience with failure experiments, disaster-recovery exercises, and validated service failovers.
  • Experience commanding SEV1 or SEV2 incidents and driving systemic remediation.
  • Experience delivering measurable reliability outcomes involving availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
  • Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
  • Strong software engineering and automation skills using Python or Go, infrastructure as code, and modern delivery toolchains.
  • Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond one team.
  • Preferred qualifications include AIOps, autonomous remediation, Amazon Bedrock AgentCore or comparable agentic frameworks, Jira and Confluence integrations, FinOps, global traffic management, progressive delivery, and global on-call operations.

Benefits

  • Base salary range of $152,000–$228,000 USD, plus bonus, equity, and benefit options.
  • Medical, dental, and vision plans; matching 401(k); unlimited PTO; paid holidays; parental/adoption leave; legal insurance; and a home technology stipend.
  • Based in the Santa Clara, California office with the option to work remotely a few days per week.
  • Up to 25% travel.

Categories

DevOpsSite Reliability
IonQ

About IonQ

1,001-5,000 employees

IonQ builds trapped-ion quantum computers and offers access to them via major cloud platforms and direct systems for enterprise and research users. The public company (NYSE: IONQ), founded in 2015 and headquartered in College Park, Maryland, sells quantum hardware, cloud services, and development tools used in areas like materials modeling, optimization, and machine learning. Its systems are available through AWS Braket, Microsoft Azure Quantum, and Google Cloud Marketplace.

Contact me