Cerebras Systems

ML Systems Performance Engineer

Cerebras Systems
Apply
3 months ago
Bengaluru, IndiaMid Level

Responsibilities

  • Build kernel-level and end-to-end performance models to estimate the performance of advanced and customer ML models.
  • Optimize and debug kernel microcode and compiler algorithms to improve inference speed, throughput, and compute utilization on the Cerebras WSE.
  • Debug and analyze runtime performance across systems and compute clusters.
  • Develop tooling and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster.

Requirements

  • Bachelor's, master's, or PhD in Electrical Engineering or Computer Science.
  • Strong background in computer architecture.
  • Understanding of low-level deep learning and LLM mathematics.
  • At least 3 years of experience in a relevant area such as computer architecture, CPU/GPU performance, kernel optimization, or HPC.
  • Experience working with CPU/GPU simulators.
  • Exposure to performance profiling and debugging on a system pipeline.
  • Comfort with C++ and Python.
  • Strong analytical and problem-solving abilities.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunity to publish and open source cutting-edge AI research.
  • Opportunity to work on one of the world's fastest AI supercomputers.
  • Job stability with startup vitality.
  • Non-corporate work culture that respects individual beliefs.
  • Continuous learning, growth, and team support.

Tech Stack

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me