Cerebras Systems

Machine Learning Engineer (Model Bring-Up)

Cerebras Systems
Apply
4 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Bring up new models by understanding architectures, loading and converting weights, implementing supported execution paths, and validating correctness against reference implementations.
  • Develop and extend MLIR dialects, graph transformations, rewrite patterns, lowering passes, and hardware-specific mappings.
  • Enable model operations including attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel changes.
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
  • Tune inference prefill and decode performance, KV-cache management, batching, and quantization for latency, throughput, and memory efficiency.
  • Investigate numerical differences and measure the accuracy impact of precision changes and compiler optimizations.
  • Use profiling, execution traces, and hardware counters to diagnose compute, memory, communication, and runtime bottlenecks.
  • Collaborate with hardware, compiler, kernel, and runtime teams to deliver reliable model support and repeatable performance benchmarks.

Requirements

  • Strong programming skills in C++ and Python.
  • Hands-on experience bringing up and debugging machine learning models in PyTorch or a comparable framework.
  • Practical experience with MLIR, including dialects, rewrite patterns, transformation passes, and lowering pipelines.
  • Understanding of compiler fundamentals including intermediate representations, dataflow analysis, and code generation.
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Ability to debug correctness and performance issues across model code, compiler-generated code, kernels, and runtime execution.
  • Preferred qualifications include experience with LLM inference, GQA, sliding-window attention, MoE, KV caching, speculative decoding, FP16, BF16, FP8, low-bit quantization, accelerator kernels, hardware-specific compiler backends, distributed execution, model parallelism, accelerator memory hierarchies, and contributions to MLIR, LLVM, inference frameworks, or related open-source projects.

Benefits

  • Opportunity to build an AI platform beyond GPU constraints and work on a high-performance AI supercomputer.
  • Opportunities to publish and open-source AI research.
  • Job stability with startup vitality and a non-corporate work culture.
  • The company describes a work environment supporting inclusion, continuous learning, growth, and employee support.

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me