
Cluster Operations Software Engineer
Cerebras Systems27 days ago
Responsibilities
- Deploy, configure, and debug container-based services using Docker.
- Build monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling for cluster operations.
- Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management.
- Manage and operate multiple advanced AI compute infrastructure clusters.
- Monitor cluster health, troubleshoot issues, and maximize compute capacity through optimization and resource allocation.
- Provide 24/7 monitoring and hands-on technical support while handling engineering escalations.
Requirements
- 6–8 years of relevant experience managing and operating complex compute infrastructure, preferably for machine learning or high-performance computing.
- Proficiency in Python and Go, including experience building operational platforms, workflow automation systems, and reliability tooling.
- Expertise in distributed systems and Linux-based compute systems and command-line tools.
- Extensive knowledge of Docker containers and container orchestration platforms such as Kubernetes.
- Experience with monitoring and alerting systems and resolving complex technical issues.
- Willingness to participate in a 24/7 on-call rotation.
- Preferred qualifications include experience operating large-scale AI clusters and knowledge of Ethernet, RoCE, TCP/IP, AWS, GCP, or Azure.
- Strong ownership, communication, collaboration, and problem-solving abilities.
Benefits
- Work with Cerebras’s Wafer-Scale Engine and one of the fastest AI supercomputers.
- Opportunities to build an AI platform beyond GPU constraints and contribute to cutting-edge AI research and open source projects.
- Job stability with startup vitality and a non-corporate work culture.
- Work location options include the San Francisco Bay Area, Toronto, Canada, or Bangalore, India.
- Participation in a 24/7 on-call rotation is required.
Categories
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.