
Principal SRE - AI Inference
Cerebras Systems2 months ago
Responsibilities
- Define and implement strategies for delivering and operating software reliably and at scale across multiple datacenters and cloud-based environments.
- Architect self-service platforms and internal tooling for safely triggering and observing critical workflows.
- Define reliability practices for inference workloads, including SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting.
- Drive architecture for capacity management, workload placement, rollout safety, validation, fleet management, and operational automation.
- Mentor senior SREs, support critical incident escalations, and prioritize automation based on production pain points.
- Measure impact through toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Requirements
- 15+ years of experience in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at large scale.
- Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.
- Experience driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms.
- Strong judgment in consolidating fragmented workflows, tools, and teams into coherent architectures.
- Ability to lead complex technical programs end to end, influence senior cross-functional stakeholders, mentor senior engineers, and communicate technical strategy.
- Hands-on production experience with observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and operational review loops.
- Preferred experience with Bazel or other large-scale build systems.
- Preferred background in AI/ML inference systems, model serving runtimes, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.
- Preferred experience with predictive autoscaling, chaos engineering, or cost-aware capacity management for compute-intensive workloads.
Benefits
- The role does not require 24/7 on-call rotations.
- The position is located in the SF Bay Area or Toronto.
- Employees can work on a breakthrough AI platform, cutting-edge AI research, and one of the world’s fastest AI supercomputers.
- Cerebras offers job stability with startup vitality and a non-corporate culture focused on learning, growth, and support.
- Cerebras is committed to an equal and diverse work environment.
Tech Stack
Bazel
Categories
Site Reliability
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.