24 hours ago
Remote, United States +2 moreStaff+
Base Salary
$132k - $211k/yr
Responsibilities
- Lead, mentor, and technically develop the Site Reliability Engineering team across multiple locations.
- Set the team’s technical direction, priorities, reliability roadmap, and standards.
- Design and maintain a sustainable on-call rotation while monitoring page load and team health.
- Define and govern SLI, SLO, and SLA frameworks for contracted availability targets up to 99.95%.
- Serve as the Tier 2 technical escalation point for major incidents and lead systemic improvements from root-cause analysis.
- Champion blameless postmortems and continuously improve Production Readiness and NFR reviews.
- Approve high-risk and out-of-window production changes.
- Set strategic direction for metrics, dashboards, alerting, and operational automation.
- Drive CI/CD automation for service deployments, rollbacks, and operational tasks.
- Partner with DevOps, platform, development, and architecture teams to embed reliability into the SDLC and evolve shared infrastructure.
- Participate in reliability consulting and architectural reviews and communicate reliability posture and risk to technical and non-technical stakeholders.
Requirements
- 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function.
- Demonstrated experience setting technical direction and holding standards across a team, with or without formal authority.
- Hands-on experience with Kubernetes, Docker, and Istio.
- Experience with Azure, AWS, and Google Cloud, with Azure as the primary platform.
- Familiarity with Zabbix, Prometheus, Grafana, and other observability tooling.
- Experience with CI/CD pipelines and infrastructure-as-code practices using tools such as Terraform and Flux.
- Proficiency in at least one of Python, Go, or Shell.
- Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals involving Layer 4/5, DNS, HTTP/S, and TLS.
- Excellent written and verbal communication skills in English.
- Preferred: site reliability leadership, distributed or multi-site technical team leadership, high-availability service design, Loki or Thanos, Jira or Confluence, and automotive, embedded, or latency-sensitive production experience.
Benefits
- Annual bonus opportunity.
- Medical, dental, vision, life, and disability insurance coverage.
- Paid time off and paid holidays.
- Company contribution to a 401(k) plan or RRSP in Canada.
- Equity awards for certain positions and levels.
- Remote and/or hybrid work is available depending on the position.
Categories
Site Reliability
About Cerence
Cerence builds automotive voice and conversational AI—speech recognition, natural language understanding, text-to-speech, and generative assistants—embedded in vehicles for automakers and transportation OEMs. Spun out of Nuance in 2019 and headquartered in Burlington, Massachusetts, the public company (NASDAQ: CRNC) licenses software and services for in-car infotainment and HMI. Its technology ships in over 500 million vehicles, with customers including brands such as BMW, Audi, Ford, and Daimler.