
Staff Software Engineer, Inference API
Cerebras Systems3 hours ago
Toronto, CanadaStaff+
Responsibilities
- Build and maintain production APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs.
- Create consistent serving semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.
- Integrate foundation models, tokenizers, prompt formats, sampling methods, attention variants, and model-specific capabilities.
- Own API compatibility, versioning, validation, deprecation, and backward-compatibility practices.
- Integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components.
- Build request routing, state transfer, error handling, retries, and lifecycle management for disaggregated inference.
- Optimize latency, time to first token, throughput, streaming, batching, serialization, tokenization, scheduling, and component communication.
- Develop correctness validation, conformance tests, workload replay, model-validation suites, benchmarks, integration tests, and release gates.
- Build observability and developer-experience tooling, including logging, tracing, metrics, dashboards, health checks, SDKs, documentation, and debugging workflows.
- Collaborate with compiler, runtime, kernel, cloud, product, solutions, and customer-facing teams to deliver scalable serving capabilities.
Requirements
- 5+ years of software engineering experience with substantial individual-contributor ownership of production software or distributed systems.
- Strong Python programming ability and experience with performance-sensitive or highly concurrent services in C++, Go, or a similar systems language.
- Hands-on experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform.
- Understanding of LLM inference concepts including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration.
- Experience integrating software across service, framework, runtime, and infrastructure boundaries and building stable APIs with validation, error handling, observability, compatibility, and versioning.
- Experience with Linux, containers, Kubernetes or comparable orchestration systems, and operating latency-sensitive production services.
- Bachelor’s degree in computer science, computer engineering, electrical engineering, or a related discipline, or equivalent practical experience.
- Ability to diagnose correctness, reliability, and performance issues across distributed serving systems and communicate effectively across functions.
- Preferred experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs; contributions to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or other open-source ML systems.
- Preferred experience with transformer, Mixture-of-Experts, diffusion, embedding, reranking, multimodal, disaggregated prefill/decode, KV-cache transfer, prefix caching, chunked prefill, admission control, or request scheduling systems.
Benefits
- Opportunity to build a breakthrough AI platform beyond GPU constraints and work on a high-performance AI supercomputer.
- Opportunities to publish and open-source AI research.
- Job stability with startup vitality and a simple, non-corporate work culture.
- Cerebras states that it is an equal opportunity employer committed to an inclusive environment with continuous learning, growth, and support.
Categories
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.