Figure AI

Helix AI Engineer, Training Performance

Figure AI
Apply
23 days ago
San Jose, CA, USASenior

Base Salary

$200k - $400k/yr

Responsibilities

  • Optimize training performance for 100B+ parameter models across 100,000+ GPUs.
  • Advise on accelerator selection, cluster topology, scheduling, networking, and hardware procurement for future scaling.
  • Write and optimize custom Triton and CUDA kernels.
  • Build tooling and dashboards for performance monitoring, regression detection, and root-cause analysis.
  • Optimize data loading and preprocessing pipelines to prevent I/O bottlenecks.
  • Improve checkpointing, fault tolerance, and elastic restart for large-scale training jobs.
  • Co-design model architectures and training recipes with researchers, including activation checkpointing, mixed precision, and sequence packing strategies.
  • Extend kernel compilers such as Triton and Gluon for faster iteration and support for custom and non-NVIDIA accelerators.
  • Build agentic systems that generate, benchmark, and iterate on custom kernels.
  • Evaluate emerging accelerator architectures and lead proof-of-concept ports and benchmarks.
  • Evaluate model and data parallelism strategies including FSDP, context parallelism, and expert parallelism.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • 3+ years of experience in AI performance engineering, including significant experience leading large-scale performance improvement projects.
  • Deep understanding of GPU architecture and performance characteristics, including memory bandwidth, compute-bound versus memory-bound operations, and occupancy.
  • Proficiency with profiling tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, HTA, or similar tools.
  • Understanding of collective communication and modern networking concepts including NCCL, RDMA, NVLink, InfiniBand/RoCE, and topology-aware placement.
  • Strong Python and CUDA/C++ skills, with the ability to read and modify framework internals.
  • Experience debugging performance regressions and instability at scale, including stragglers, hangs, out-of-memory issues, and numerical divergence.
  • Experience defining and using hardware-efficiency metrics such as MFU and HFU.
  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration is a bonus.
  • Open-source contributions to ML systems projects such as PyTorch, Megatron-LM, vLLM, DeepSpeed, or JAX are a bonus.
  • Exposure to non-NVIDIA accelerators and heterogeneous fleet management is a bonus.

Benefits

  • This is a full-time position.
  • Additional compensation components and benefits may be provided depending on the specific role and shared if an employment offer is extended.

Categories

Contact me