Cerebras Systems

Full Stack LLM Engineer

Cerebras Systems
Apply
1 year ago
Toronto, CanadaSenior

Responsibilities

  • Contribute to the end-to-end bringup of ML models on Cerebras CSX systems.
  • Work across model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
  • Debug performance and correctness issues spanning model code, compiler IRs, runtime behavior, and hardware utilization.
  • Prototype improvements to tools, APIs, and automation flows to accelerate future model bringups.

Requirements

  • Bachelor’s, master’s, or PhD in computer science, engineering, or a related field.
  • Comfort navigating Python modeling code, compiler IRs, performance profiling, and the broader AI toolchain.
  • Strong debugging skills across performance, numerical accuracy, and runtime integration.
  • Experience with deep learning frameworks such as PyTorch and TensorFlow, plus familiarity with attention, mixture-of-experts, and diffusion model internals.
  • Proficiency in C/C++ programming and experience with low-level optimization.
  • Proven compiler development experience, particularly with LLVM and/or MLIR.
  • Strong background in optimization techniques, particularly those involving NP-hard problems.

Benefits

  • Competitive salary and benefits package.
  • Professional growth and career advancement opportunities.
  • Dynamic and innovative work environment.
  • Opportunity to work on cutting-edge AI technologies and contribute to the future of AI.
  • Opportunity to work on an AI platform beyond GPU constraints, publish and open-source AI research, and work on a high-performance AI supercomputer.
  • Job stability with startup vitality and a non-corporate work culture.

Tech Stack

CC++PythonPyTorchTensorFlow

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me