
Staff Site Reliability Engineer – Automation and Platform
Cerebras Systems11 months ago
Responsibilities
- Define and implement strategies for reliably delivering and running software at scale across multiple datacenters and cloud-based environments.
- Architect self-service platforms and internal tooling for safely triggering and observing critical workflows.
- Develop reliability practices for inference workloads, including SLOs, SLIs, error budgets, blameless postmortems, chaos testing, and capacity forecasting.
- Automate operational toil and use production pain points to prioritize high-leverage engineering work.
- Mentor SREs, support critical incident escalations, and collaborate with core, cluster, cloud, and product stakeholders.
- Measure impact through toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Requirements
- 8+ years of experience in SRE, infrastructure engineering, or platform engineering.
- Strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments.
- Deep expertise operating large-scale heterogeneous clusters with a proprietary cloud control plane.
- Experience designing and delivering CI/CD or GitOps systems using Argo CD or similar tools, with strong safety and observability.
- Hands-on experience with observability systems such as Loki, Tempo, Mimir, and Prometheus.
- Ability to lead complex projects end to end, influence cross-functional stakeholders, and communicate technical direction clearly.
- Experience with Bazel or other large-scale build systems in production is preferred.
- Experience with AI/ML inference systems, model-serving runtimes, GPU or wafer-scale orchestration, latency and accuracy SLOs, or drift monitoring is preferred.
- Experience with predictive autoscaling, chaos engineering, or cost-aware capacity planning for compute-intensive workloads is preferred.
Benefits
- Work on a breakthrough AI platform and high-performance wafer-scale AI supercomputer.
- Opportunities to publish and open-source cutting-edge AI research.
- Startup vitality with job stability and a simple, non-corporate work culture.
- Locations are the SF Bay Area or Toronto.
- The role does not require 24/7 on-call rotations.
Tech Stack
Argo CDBazelPrometheus
Categories
DevOpsSite Reliability
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.