
AI Fleet Platform Software Engineer
Cerebras Systems10 hours ago
Responsibilities
- Build and operate software for managing large fleets of AI clusters.
- Develop operator-facing views of cluster health, capacity, performance, and ongoing issues.
- Build services and integrations that combine data and workflows from multiple infrastructure systems.
- Automate repetitive operational work and improve incident investigation and service restoration tools.
- Design reliable systems that continue functioning through component and site failures.
- Partner with platform users and cross-functional infrastructure and operations teams to make practical engineering decisions.
- Lead projects from initial design through production, measure their impact, and improve them using operational feedback.
Requirements
- 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure.
- Strong Go or Python skills and experience designing services and APIs.
- Expertise in control planes, fleet management systems, or operational platforms.
- Experience with Linux, containers, Kubernetes, and failure handling in distributed systems.
- Experience designing systems with asynchronous work, retries, and partial failures.
- Experience with event streaming, workflow automation, or time-series telemetry.
- Strong judgment in reliability, security, and observability.
- Ability to lead ambiguous projects and collaborate effectively across engineering and operations teams.
- Preferred experience includes incident response, hardware health, capacity management, infrastructure dashboards, and AI cluster compute, networking, and hardware systems.
Benefits
- The role is based in the San Francisco Bay Area or Toronto, Canada.
- Cerebras highlights opportunities to work on an AI platform beyond GPU constraints, publish and open-source AI research, work on a fast AI supercomputer, and participate in a startup environment with job stability.
- The company describes a non-corporate culture that respects individual beliefs and supports learning, growth, and inclusion.
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.