
ML Algorithm Mapping and Performance Engineer, Core ML
Cerebras Systems16 hours ago
Responsibilities
- Build analytical and empirical performance models for ML training and inference algorithms.
- Analyze algorithmic scaling and trade-offs across model size, sequence length, batch size, parallelism, and hardware scale.
- Construct Pareto frontiers covering model quality, latency, throughput, memory, communication, and compute cost.
- Develop prototype implementations and benchmarks for the Cerebras WSE and GPU or software baselines.
- Identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.
- Evaluate parallel token generation, diffusion, speculative decoding, attention, sparsity, mixture-of-experts, low-precision computation, and distributed training techniques.
- Collaborate with research, kernel, compiler, runtime, inference, and architecture teams on implementation and co-design directions.
- Develop tools and visualizations for performance projections, measurements, and design trade-offs.
- Communicate findings, assumptions, limitations, and recommendations through reports, presentations, and design reviews.
Requirements
- Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field.
- Strong foundation in computer architecture, parallel computing, and systems performance.
- Strong understanding of machine learning fundamentals and ML systems, including compute, memory, communication, accuracy, and scaling behavior.
- Experience with analytical performance modeling, algorithmic complexity analysis, benchmarking, or system simulation.
- Strong ability to reason from first principles about compute, memory, and communication costs.
- Proficiency in Python and comfort with C++.
- Experience profiling and debugging performance in ML, HPC, CPU, GPU, or accelerator-based systems.
- Ability to combine mathematical analysis, experimental validation, and practical engineering recommendations.
- Preferred experience with roofline analysis, CPU or GPU simulators, kernel optimization, or hardware-software co-design.
- Preferred familiarity with CUDA, Triton, PyTorch, JAX, or open-source LLM training and inference systems.
- Preferred understanding of transformer internals, attention variants, KV-cache strategies, model parallelism, sparsity, quantization, and parallel generation.
- Research publications, patents, significant open-source contributions, or experience evaluating future hardware and software architectures are preferred.
Benefits
- Opportunity to work on Cerebras’s wafer-scale AI platform and AI supercomputers.
- Opportunities to publish and open-source cutting-edge AI research.
- Job stability with startup vitality and a non-corporate work culture.
- Equal opportunity workplace focused on inclusion, learning, growth, and support.
Categories
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.