
MTS, Research Engineer
Fireworks AIabout 1 month ago
Base Salary
$250k - $400k/yr
Responsibilities
- Explore new model architectures, training objectives, and optimization techniques through hypothesis-driven experiments.
- Implement and reproduce results from recent machine learning papers, identify limitations, and scale methods to larger datasets and models.
- Design, implement, maintain, and optimize high-performance distributed machine learning systems across large GPU clusters.
- Optimize training loops, data loaders, and communication overhead.
- Translate mathematical concepts and research ideas into robust, efficient, and maintainable code.
- Collaborate with Research Scientists by providing tooling, optimizing code, and co-designing hardware-aware experiments.
Requirements
- Strong programming skills in Python, C++, or Rust and a commitment to clean, maintainable code.
- Deep practical knowledge of PyTorch, JAX, or TensorFlow.
- Experience with large distributed systems and parallel computing, including technologies such as CUDA, NCCL, or MPI.
- Strong foundation in linear algebra, calculus, probability, and statistics.
- Proven track record of implementing complex deep learning algorithms from scratch.
- Master’s or PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related field, or equivalent industry experience, is preferred.
- Experience with low-level GPU programming using CUDA or Triton, or with hardware co-design, is preferred.
- Familiarity with large language model training, inference, SGLang, and vLLM is preferred.
Benefits
- Equal-opportunity and inclusive workplace
- Opportunity to work on cutting-edge AI infrastructure and model serving
- Collaboration with world-class engineers and AI researchers