
Staff Software Engineer, Inference API
Cerebras Systems3 hours ago
Toronto, CanadaStaff+
Responsibilities
- Build and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs.
- Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.
- Integrate foundation models, tokenizers, prompt formats, sampling methods, attention variants, and model-specific capabilities into the serving platform.
- Own API compatibility, versioning, deprecation, validation, and backward compatibility practices.
- Integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components.
- Build request routing, state transfer, error handling, retries, and lifecycle management for disaggregated inference.
- Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components.
- Develop correctness validation, conformance tests, model-validation suites, performance benchmarks, integration tests, workload-replay tools, and release gates.
- Strengthen production reliability and observability through logging, tracing, metrics, dashboards, health checks, and diagnostic tooling.
- Build SDKs, documentation, examples, debugging tools, configuration, and self-service workflows for developers, customers, and partners.
- Collaborate with compiler, runtime, kernel, cloud, product, solutions, and customer-facing teams.
Requirements
- At least 5 years of software engineering experience with substantial individual-contributor ownership of production software or distributed systems.
- Strong programming ability in Python and Go, plus experience with C++, Rust, or a similar systems language for performance-sensitive or highly concurrent services.
- Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices.
- Experience integrating software across service, framework, runtime, and infrastructure boundaries.
- Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs.
- Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and latency-sensitive production services.
- Ability to diagnose correctness, reliability, and performance issues across distributed serving systems.
- Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
- Preferred experience contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project.
- Preferred experience with API conformance, model-quality, numerical-comparison, determinism, performance-regression, SDK, developer-tool, model-registry, or self-service ML platform development.
- Preferred understanding of multi-model or multi-tenant inference, routing, admission control, fairness, quotas, rate limiting, capacity-aware scheduling, tokenization, generation configuration, logits processing, structured generation, and constrained decoding.
- Preferred experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, request scheduling, reduced-precision inference, and quantization formats such as BF16, FP8, FP4, INT8, or INT4.
Benefits
- Opportunity to build an AI platform beyond GPU constraints and work on one of the world's fastest AI supercomputers.
- Opportunities to publish and open source cutting-edge AI research.
- Job stability with startup vitality and a simple, non-corporate work culture.
- Inclusive equal-opportunity work environment focused on learning, growth, and support.
Categories
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.