Cerebras Systems

Staff GPU Inference SDET

Cerebras Systems
Apply
4 hours ago
Toronto, Canada or Sunnyvale, CA, USAStaff+
H1B sponsor

Responsibilities

  • Design and implement automated test frameworks, regression gates, and release qualification pipelines for the full GPU inference stack.
  • Bring up, validate, and stress-test multi-node GPU clusters and distributed LLM serving frameworks.
  • Benchmark prefill and decode workers, continuous batching, prefix caching, KV-cache efficiency, tensor parallelism, and expert parallelism.
  • Build workload replay and performance-model verification tools tracking TTFT, ITL, throughput, P99 tail latency, and capacity efficiency.
  • Validate model accuracy, precision stability, determinism, and output correctness across software updates, kernel fusions, and hardware revisions.
  • Develop chaos-engineering and fault-injection suites for node failures, network degradation, memory leaks, driver or firmware mismatches, and automated recovery.
  • Integrate automated testing with telemetry and continuous performance monitoring systems.
  • Perform root-cause analysis and debugging across software and hardware boundaries.

Requirements

  • 8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
  • Hands-on experience provisioning and validating multi-node NVIDIA or AMD GPU clusters in public cloud or enterprise data center environments.
  • Deep understanding of LLM serving engines, distributed runtimes, prefill/decode disaggregation, KV-cache management, and dynamic batching.
  • Expert-level Python skills and extensive experience building test automation frameworks, diagnostic tooling, and CI/CD integrations.
  • Strong proficiency with Kubernetes, Slurm, or Ray and high-performance interconnects such as InfiniBand, RoCE, or NCCL.
  • Proven experience with root-cause analysis, stress testing, and node-failure simulation in distributed systems.
  • Preferred experience with AMD ROCm/HIP or NVIDIA software stacks.
  • Preferred experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.
  • Familiarity with PyTorch Profiler, NVTX, ROCm profilers, or C++ is preferred.

Benefits

  • Opportunity to build a breakthrough AI platform and work on a high-performance AI supercomputer.
  • Opportunities to publish and open-source AI research.
  • Job stability with startup vitality and a non-corporate work culture.
  • Cerebras states that it supports continuous learning, growth, and an inclusive, equal-opportunity work environment.

Tech Stack

C++GrafanaKubernetesPrometheusPython

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me