Cerebras Systems

Staff Software Engineer, Inference Cloud

Cerebras Systems
Apply
2 years ago
Sunnyvale, CA, USAStaff+
H1B sponsor

Responsibilities

  • Shape the technical direction and roadmap for major areas of the Inference Cloud Platform.
  • Design and build service discovery, request routing, load balancing, caching, batching, and traffic-management components for AI inference workloads.
  • Architect active-active, multi-region systems with rapid failover, graceful degradation, high availability, and low latency.
  • Define admission control, quota management, rate limiting, and differentiated quality-of-service mechanisms.
  • Write and review production code and make high-consequence architectural decisions within the platform domain.
  • Lead complex production issues, improve observability and incident response, and drive capacity planning and post-incident improvements.
  • Partner with ML, Product, Infrastructure, and Platform teams on scalable system designs and shared technical decisions.
  • Mentor senior engineers through design feedback, pairing, and technical standards.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep expertise in distributed-systems architecture in cloud environments, including networking, compute orchestration, container platforms, and multi-region production services.
  • Strong record of architectural decision-making for highly available, latency-sensitive systems at scale.
  • Experience optimizing latency, throughput, and efficiency in high-QPS systems; experience reducing TTFT and tail latency is a strong plus.
  • Strong proficiency in backend or systems languages such as Go, C++, or Python, with the ability to contribute production code directly.
  • Experience designing observability and reliability practices involving metrics, logging, tracing, alerting, incident response, and SLO-driven operations.
  • Ability to influence senior engineers and cross-functional partners through technical credibility, communication, and judgment.
  • Preferred experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunity to publish and open source cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Job stability with startup vitality.
  • Simple, non-corporate work culture that respects individual beliefs.
  • Equal and diverse work environment with continuous learning, growth, and support.

Tech Stack

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me