Applied Intuition

Machine Learning Performance Engineer - Offboard Training & Inference

Applied Intuition
Apply
27 days ago
Sunnyvale, CA, USASenior
H1B sponsor

Base Salary

$215k - $285k/yr

Responsibilities

  • Profile and optimize distributed training across data loading, preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
  • Optimize large-scale offline and batch inference through batching, scheduling, quantization, low-precision execution, graph optimization, and accelerator saturation.
  • Develop roofline and performance models, quantify achieved versus theoretical performance, and prioritize optimization opportunities.
  • Improve multi-node scaling through sharding, parallelism, collective communication, interconnect utilization, memory-bandwidth optimization, and kernel fusion.
  • Increase cluster goodput by reducing GPU idle time caused by input stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery.
  • Build benchmarking, observability, and regression-detection tooling for evolving models and code.
  • Collaborate with engineers across functions to solve large-scale data and compute problems and contribute to technical and product decisions.

Requirements

  • Hands-on production experience with ML performance engineering, profiling, roofline analysis, throughput optimization, and root-cause investigation.
  • Experience with distributed multi-node training at scale using FSDP, DeepSpeed, Megatron, NCCL, or equivalent systems.
  • Deep familiarity with GPU or accelerator performance, including memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
  • Experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, or Ray.
  • Fluency in Python and proficiency in C++ or another systems language.
  • Strong debugging, analytical, and problem-solving skills, plus a deep understanding of machine-learning foundations.
  • GPU kernel development experience with CUDA, Triton, CUTLASS, or hand-tuned attention implementations is preferred.
  • Experience with Nsight Systems, Nsight Compute, PyTorch Profiler, or perf is preferred.
  • Experience with GPU scheduling and orchestration using Kubernetes, Slurm, or Ray is preferred.
  • Experience with fault tolerance and elastic training, including checkpointing, straggler mitigation, and preemption recovery, is preferred.
  • Familiarity with autonomy or robotics data, ROS, OpenCV, and multi-sensor log formats is preferred.

Benefits

  • Primarily in-office work five days per week, with occasional remote-work flexibility and schedule flexibility for personal commitments.
Applied Intuition

About Applied Intuition

1,001-5,000 employees

Applied Intuition builds software tools, infrastructure, and operating systems for developing, testing, and deploying autonomous and driver-assistance systems across automotive, trucking, defense, construction, mining, and agriculture. It sells enterprise software and services including simulation, sensor data management, HD mapping, and a vehicle OS to OEMs and government customers. Founded in 2017 and headquartered in Sunnyvale, California, it counts 18 of the top 20 global automakers and the U.S. military among its users.

Contact me