Cerebras Systems

Sr. Staff Software Engineer, Inference Platform

Cerebras Systems
Apply
3 months ago
Toronto, Canada or Sunnyvale, CA, USAStaff+
H1B sponsor

Responsibilities

  • Design, develop, test, maintain, and productionize software for the inference platform.
  • Shape platform direction, including Kubernetes custom resource definitions, failure domains, service boundaries, system evolution, and roadmaps for major technical areas.
  • Architect active-active systems with rapid failover, graceful degradation, clear SLOs, and improved latency, throughput, capacity efficiency, and resilience.
  • Write and review production code, make high-consequence architectural decisions, and set technical standards through design and code reviews.
  • Lead difficult production issues and cross-system bottlenecks, including observability, incident response, capacity planning, and post-incident improvements.
  • Raise the effectiveness of senior engineers through design feedback, pairing, and technical standards.
  • Partner with ML, Product, Infrastructure, and Cloud teams on scalable system designs and shared technical decisions.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor work building and operating large-scale distributed systems or cloud infrastructure.
  • Deep expertise in distributed-systems architecture, ideally with Kubernetes.
  • Strong experience making architectural decisions for highly available, latency-sensitive systems at scale.
  • Experience with security certificates, TLS, and mTLS.
  • Experience optimizing latency, throughput, and efficiency in high-QPS systems; TTFT and tail-latency reduction experience is a strong plus.
  • Strong proficiency in backend or systems languages such as Go or C++, with the ability to contribute production code directly.
  • Experience designing observability and reliability practices involving metrics, logging, tracing, alerting, incident response, and SLO-driven operations.
  • Ability to influence senior engineers and cross-functional partners through technical credibility, communication, and judgment.
  • Preferred: experience with ML inference infrastructure, model-serving systems, or GPU-accelerated workloads.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunities to publish and open source cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Job stability with startup vitality and a simple, non-corporate work culture.
  • Sunnyvale or Toronto preferred.
  • Cerebras is committed to an equal and diverse workplace and supports continuous learning, growth, and development.

Tech Stack

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me