Cerebras Systems

Principal Engineer, Inference Cloud

Cerebras Systems
Apply
11 months ago
Sunnyvale, CA, USAStaff+
H1B sponsor

Responsibilities

  • Identify and prioritize high-leverage platform problems and make explicit architectural and product-support tradeoffs.
  • Set the long-term direction for multi-region topology, failure domains, service boundaries, and platform evolution.
  • Architect active-active systems with failover, circuit breaking, backpressure, load shedding, and clear SLOs.
  • Improve latency, throughput, capacity efficiency, and resilience under unpredictable, bursty AI workloads.
  • Write production code on critical paths and review designs and implementations.
  • Lead complex production issues, incident response, observability, capacity planning, and post-incident improvements.
  • Drive cross-team decisions involving reliability, API design, capacity planning, deployment strategy, and shared infrastructure.
  • Mentor engineers and improve technical decision-making through design feedback, pairing, and engineering standards.

Requirements

  • 10+ years of software engineering experience, including substantial individual-contributor experience with large-scale distributed systems or cloud infrastructure.
  • Deep expertise in distributed-systems architecture in cloud environments, networking, compute orchestration, container platforms, and multi-region production services.
  • Demonstrated experience making architectural decisions for highly available, latency-sensitive systems at scale.
  • Experience optimizing latency, throughput, and efficiency in high-QPS systems; TTFT and tail-latency reduction experience is a strong plus.
  • Strong proficiency in Go, C++, or Python and the ability to contribute production code directly.
  • Experience designing observability and reliability practices involving metrics, logging, tracing, alerting, incident response, and SLI/SLO/SLA-driven operations.
  • Ability to influence senior engineers, technical leads, and cross-functional partners through technical credibility, communication, and judgment.
  • Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads is preferred.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunity to publish and open-source cutting-edge AI research.
  • Opportunity to work on one of the fastest AI supercomputers in the world.
  • Job stability with startup vitality.
  • A simple, non-corporate work culture that respects individual beliefs.
  • Equal and diverse work environment with continuous learning, growth, and support.

Tech Stack

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me