Cerebras Systems

Staff Software Engineer, GPU Inference

Cerebras Systems
Apply
1 month ago
Toronto, Canada or Sunnyvale, CA, USAStaff+
H1B sponsor

Responsibilities

  • Design, build, deploy, and maintain the complete GPU prefill path across APIs, serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, compatibility, and failure-recovery practices for the AMD GPU fleet.
  • Define service-level indicators and objectives and improve fault isolation, graceful degradation, automated recovery, incident response, and remediation.
  • Profile and optimize time to first token, throughput, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
  • Improve scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
  • Debug failures and performance regressions across application code, serving runtimes, frameworks, communication libraries, kernels, drivers, firmware, networking, and hardware.
  • Build validation and regression systems for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
  • Develop benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
  • Strong C++ and Python programming skills, including multithreading, concurrency, memory management, and performance-sensitive software.
  • Hands-on experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent internally developed system.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
  • Experience debugging distributed systems across multiple layers and operating latency-sensitive services in production.
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, and CI/CD.
  • Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
  • Strong communication and technical leadership skills, with the ability to drive ambiguous cross-functional projects to completion.
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
  • Preferred: AMD Instinct and ROCm experience, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, Composable Kernel, or related tools.
  • Preferred: CUDA, open-source ML systems contributions, disaggregated inference, multi-GPU and multi-node inference, GPU kernel optimization, reduced-precision inference, quantization, numerical validation, and performance-regression testing experience.

Benefits

  • Opportunity to build a breakthrough AI platform and work on a high-performance AI supercomputer.
  • Opportunities to publish and open source cutting-edge AI research.
  • Job stability with startup vitality and a simple, non-corporate culture.
  • Inclusive equal-opportunity work environment focused on continuous learning, growth, and support.
Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me