Cerebras Systems

Cluster Operations Software Engineer

Cerebras Systems
Apply
27 days ago
Toronto, Canada +2 moreSenior
H1B sponsor

Responsibilities

  • Deploy, configure, and debug container-based services using Docker.
  • Build monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling for cluster operations.
  • Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management.
  • Manage and operate multiple advanced AI compute infrastructure clusters.
  • Monitor cluster health, troubleshoot issues, and maximize compute capacity through optimization and resource allocation.
  • Provide 24/7 monitoring and hands-on technical support while handling engineering escalations.

Requirements

  • 6–8 years of relevant experience managing and operating complex compute infrastructure, preferably for machine learning or high-performance computing.
  • Proficiency in Python and Go, including experience building operational platforms, workflow automation systems, and reliability tooling.
  • Expertise in distributed systems and Linux-based compute systems and command-line tools.
  • Extensive knowledge of Docker containers and container orchestration platforms such as Kubernetes.
  • Experience with monitoring and alerting systems and resolving complex technical issues.
  • Willingness to participate in a 24/7 on-call rotation.
  • Preferred qualifications include experience operating large-scale AI clusters and knowledge of Ethernet, RoCE, TCP/IP, AWS, GCP, or Azure.
  • Strong ownership, communication, collaboration, and problem-solving abilities.

Benefits

  • Work with Cerebras’s Wafer-Scale Engine and one of the fastest AI supercomputers.
  • Opportunities to build an AI platform beyond GPU constraints and contribute to cutting-edge AI research and open source projects.
  • Job stability with startup vitality and a non-corporate work culture.
  • Work location options include the San Francisco Bay Area, Toronto, Canada, or Bangalore, India.
  • Participation in a 24/7 on-call rotation is required.
Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me