Applied Intuition

Machine Learning Performance Engineer - Offboard Training & Inference

Applied Intuition
Apply
about 5 hours ago
Sunnyvale, CA, USASenior
H1B Sponsor

Base Salary

$215k - $285k/yr

Responsibilities

  • Profile and optimize distributed training end to end, including data loading, preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs through batching, scheduling, quantization, low-precision execution, graph optimization, and accelerator saturation.
  • Establish roofline and performance models, measure the gap between achieved and theoretical performance, and prioritize optimization opportunities by impact and effort.
  • Improve multi-node scaling through sharding, parallelism, collective communication, interconnect utilization, memory-bandwidth optimization, and kernel fusion.
  • Increase cluster goodput by reducing GPU idle time caused by input stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery.
  • Build benchmarking, observability, and regression-detection tooling to prevent performance degradation as models and code evolve.
  • Collaborate with engineers across functions to solve complex data and compute problems at scale.
  • Contribute to a collaborative, technically excellent, and innovative engineering culture.

Requirements

  • Hands-on production experience with ML performance engineering, profiling, roofline analysis, throughput optimization, and root-cause investigation.
  • Experience with distributed multi-node training at scale using FSDP, DeepSpeed, Megatron, NCCL, or an equivalent technology, including diagnosing scaling inefficiencies.
  • Deep familiarity with GPU or accelerator performance concepts such as memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
  • Experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, or Ray.
  • Fluency in Python and proficiency in C++ or another systems language.
  • Strong debugging, analytical, and problem-solving abilities, with deep understanding of machine learning foundations.
  • Ability to develop technical solutions for problems without an established playbook.
  • Preferred experience developing GPU kernels with CUDA, Triton, CUTLASS, or hand-tuned attention implementations.
  • Preferred experience with Nsight Systems, Nsight Compute, PyTorch Profiler, or perf.
  • Preferred experience with GPU scheduling and orchestration using Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization.
  • Preferred experience with fault tolerance and elastic training, including checkpointing, straggler mitigation, and preemption recovery.
  • Preferred familiarity with autonomy or robotics data, ROS, OpenCV, and multi-sensor log formats.

Benefits

  • Primarily in-office work 5 days per week, with occasional remote-work flexibility and schedule flexibility for personal commitments.
  • Equal opportunity employment and federal-contractor protections.

Tech Stack

Categories

Applied Intuition

About Applied Intuition

1,001-5,000 employees

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at applied.co.