
Machine Learning Performance Engineer
Jane Street2 months ago
Responsibilities
- Optimize machine learning model training and inference performance.
- Improve efficient large-scale training, low-latency real-time inference, and high-throughput research inference.
- Develop and optimize CUDA implementations using a whole-systems approach across storage, networking, host systems, and GPUs.
- Investigate throughput, goodput, cache behavior, latency, and other low-level performance characteristics.
- Support and optimize networking technologies and collective algorithms for distributed GPU clusters.
Requirements
- Understanding of modern machine learning techniques and toolsets.
- Experience and systems knowledge to debug training-run performance end to end.
- Low-level GPU knowledge including PTX, SASS, warps, cooperative groups, Tensor Cores, and memory hierarchy.
- Debugging and optimization experience with CUDA GDB, Nsight Systems, and Nsight Compute.
- Knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS.
- Understanding of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads.
- Background in InfiniBand, RoCE, GPUDirect, PXN, rail optimization, and NVLink for connecting GPU clusters.
- Understanding of collective algorithms for distributed GPU training in NCCL or MPI.
- Inventive approach and willingness to question existing approaches and tools.
- Fluency in English.
Tech Stack
Sass
Categories
About Jane Street
Jane Street is a global quantitative trading firm and liquidity provider that builds in-house software and research platforms to trade across asset classes. It makes markets and executes proprietary strategies on exchanges and electronic venues, serving institutional markets rather than individual investors. Founded in 2000 and headquartered in New York, it is privately held with offices in London, Hong Kong, Singapore, and Amsterdam.