Cerebras Systems

ML Runtime and Kernel Engineer - Core ML

Cerebras Systems
Apply
16 hours ago
Toronto, Canada or Sunnyvale, CA, USASenior
H1B sponsor

Responsibilities

  • Design and implement runtime components and high-performance kernels for novel Core ML algorithms.
  • Translate research prototypes into efficient Cerebras implementations, including reference implementations and GPU comparisons where useful.
  • Profile and debug performance across ML framework, compiler, runtime, communication, and kernel layers.
  • Optimize computation, memory movement, communication, and concurrency for large-scale training and low-latency inference.
  • Develop benchmarks, instrumentation, and automated tests for functionality, performance, and numerical correctness.
  • Collaborate with Core ML researchers and compiler, runtime, kernel, and inference engineers on end-to-end capabilities.
  • Contribute to software architecture and roadmap decisions by identifying platform limitations and improvements.

Requirements

  • Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels.
  • Strong programming skills in C++ and Python.
  • Understanding of parallel programming, memory management, concurrency, data structures, and performance optimization.
  • Ability to debug and profile complex software across multiple system layers.
  • Familiarity with machine learning architectures and frameworks such as PyTorch or JAX.
  • Ability to collaborate with researchers and translate evolving algorithmic requirements into reliable software.
  • Preferred experience with CUDA, Triton, low-level assembly, accelerator programming, or a C-like domain-specific language.
  • Preferred experience with compiler internals, distributed runtimes, custom hardware interfaces, or HPC systems.
  • Preferred understanding of ML fundamentals, LLM training or inference, attention, KV-cache management, parallel generation, and distributed execution.
  • Preferred experience in industrial or academic research environments and contributions to open-source systems, ML frameworks, compilers, or kernel libraries.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunity to publish and open source cutting-edge AI research.
  • Work on a high-performance AI supercomputer.
  • Job stability with startup vitality.
  • Non-corporate work culture that respects individual beliefs.
  • Equal and diverse work environment with continuous learning, growth, and support.
Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me