Cerence

Senior Principal AI Engineer

Cerence
Apply
1 month ago
Remote, United StatesStaff+

Responsibilities

  • Design and operate distributed training systems for large neural networks across GPU clusters.
  • Optimize multi-node and multi-GPU execution, GPU utilization, throughput, memory usage, and networking performance.
  • Build and optimize GPU cluster orchestration with Slurm, Kubernetes, Ray, and RunAI.
  • Debug distributed communication using NCCL, RDMA, InfiniBand, and NVLink.
  • Scale training with PyTorch Distributed, Megatron-LM, and DeepSpeed, including launch configurations, failure recovery, and performance tuning.
  • Apply activation checkpointing, ZeRO Stages 1–3, and offload strategies to increase model and batch scale.
  • Improve training stability and fault tolerance and partner with research and applied ML teams to productionize large-model training pipelines.

Requirements

  • Deep hands-on experience with distributed systems or ML systems.
  • Production experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Strong understanding of data, tensor, and pipeline parallelism strategies.
  • Low-level understanding of GPU communication and networking.
  • Experience with GPU orchestration tools including Slurm, Kubernetes, Ray, and RunAI.
  • Experience with NCCL, RDMA, InfiniBand, NVLink, PyTorch Distributed, Megatron-LM, DeepSpeed, activation checkpointing, and ZeRO offload techniques.
  • Experience as an ML Systems Engineer, Distributed Systems Engineer, AI Infrastructure Engineer, or HPC Engineer transitioning into ML is relevant.
  • Experience with large language models or foundation models is a strong plus.
Cerence

About Cerence

1,001-5,000 employees
Contact me