Senior Applied Research Engineer
Fundamental5 months ago
Barcelona, SpainSenior
Responsibilities
- Profile end-to-end distributed training runs to identify compute, GPU memory, and inter-GPU communication bottlenecks.
- Improve the efficiency and reliability of large-scale training jobs through architectural decisions and Triton/CUDA kernel development.
- Design and implement model scaling, parallelization, and memory optimization techniques for very large context sizes.
- Collaborate with ML researchers to diagnose architectural inefficiencies and ensure research ideas scale efficiently in practice.
- Drive productionization and serving of models, including improving inference efficiency through quantization.
- Share internal knowledge about model efficiency and optimization.
Requirements
- Strong understanding of modern machine learning architectures and large-scale training pipelines.
- Experience running distributed training jobs on multi-GPU systems.
- Advanced profiling and debugging skills across CPU, GPU, memory usage, latency, and inter-GPU communication.
- Strong programming skills in Python.
- Experience with model scaling and parallelization strategies, including tensor and pipeline parallelism.
- Preferred: familiarity with NCCL, MPI, and distributed communication primitives.
- Preferred: knowledge of PyTorch and Triton internals.
- Preferred: programming experience with C++ and CUDA.
Benefits
- Competitive compensation with salary and equity.
- Comprehensive health coverage for employees and dependents.
- Paid parental leave for all new parents, including adoptive and surrogate journeys.
- Relocation support for employees moving to one of the company's office locations.
- Mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action.