11 months ago
San Francisco, CA, USAMid Level
Responsibilities
- Optimize models for speed, efficiency, and reliability through low-level optimizations and systems design.
- Develop and optimize CUDA kernels, memory layouts, GPU scheduling, compute paths, and networking paths.
- Tune performance for large-scale machine-learning training and inference workloads.
- Work alongside research teams to productionize new model architectures.
- Improve and operate large-scale ML infrastructure.
Requirements
- Strong background in systems-level machine-learning engineering.
- Experience with CUDA, GPU kernel optimization, and performance tuning.
- Fluency in Python and at least one systems language, with C++ or Rust preferred.
- Familiarity with distributed training frameworks such as PyTorch, JAX, or DeepSpeed.
- Experience with large-scale training or inference infrastructure.
- Understanding of memory management, parallelization, and hardware-aware model optimization.
- At least two years of experience in ML infrastructure or performance-critical environments.
- Willingness to work in person from the San Francisco office in FiDi.
Benefits
- In-person work from the San Francisco office in FiDi.
