
Software Engineer, Inference Platform
Cerebras Systems3 months ago
Toronto, Canada or Sunnyvale, CA, USAMid Level
H1B sponsor
Responsibilities
- Design, develop, test, and maintain production software across observability, security, networking, debugging, and productionization.
- Shape the technical direction, architecture, Kubernetes custom resource definitions, failure domains, service boundaries, and roadmap for major Inference Platform areas.
- Architect active-active systems with rapid failover, graceful degradation, clear SLOs, and improved latency, throughput, capacity efficiency, and resilience.
- Write and review production code in critical platform areas and make high-consequence architectural decisions through design and code reviews.
- Lead complex production issues, incident response, capacity planning, observability efforts, and post-incident improvements.
- Partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs and align on technical decisions.
Requirements
- 3+ years of software engineering experience building and operating large-scale distributed systems or cloud infrastructure.
- Experience with distributed systems, ideally including Kubernetes.
- Experience building highly available, latency-sensitive systems at scale.
- Experience with security certificates, TLS, and mTLS.
- Experience optimizing latency, throughput, and efficiency in high-QPS systems; experience reducing TTFT and tail latency is a strong plus.
- Strong proficiency in backend or systems languages such as Go or C++.
- Preferred: experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.
Benefits
- Opportunity to build a breakthrough AI platform beyond GPU constraints.
- Opportunity to publish and open source AI research and work on a high-performance AI supercomputer.
- Job stability with startup vitality and a simple, non-corporate work culture.
- Location is open to Sunnyvale or Toronto.
- Equal opportunity employer committed to an inclusive and diverse work environment with continuous learning, growth, and support.
Tech Stack
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.