
Machine Learning Engineer — Training Optimization
Featherless AI8 months ago
Remote, WorldwideSenior
Responsibilities
- Optimize large-scale model training pipelines for throughput, convergence, stability, and cost.
- Improve distributed training strategies using data, model, and pipeline parallelism.
- Tune optimizers, schedulers, batch sizing, and bf16, fp16, and fp8 precision.
- Reduce training time and compute cost through profiling, bottleneck analysis, and systems-level improvements.
- Collaborate with researchers on architecture-aware training strategies.
- Build and maintain training infrastructure for checkpointing, fault tolerance, and reproducibility.
- Evaluate and integrate training techniques such as gradient checkpointing, ZeRO, FSDP, and custom kernels.
- Own training performance metrics and continuously improve them.
Requirements
- Strong experience training large neural networks, including LLMs or similarly large models.
- Hands-on experience with training optimization rather than only model usage.
- Solid understanding of backpropagation, optimization algorithms, training dynamics, and distributed systems for ML training.
- Experience with PyTorch is required.
- Comfort working close to hardware, including GPU, memory, and networking constraints.
- Ability to move between research ideas and production-ready code.
- Experience with large-scale distributed training across multiple nodes and GPUs is preferred.
- Familiarity with DeepSpeed, FSDP, Megatron, or custom training stacks is preferred.
- Experience optimizing training on AMD or NVIDIA GPUs is preferred.
- Contributions to open-source ML infrastructure or research codebases are preferred.
- Exposure to non-Transformer architectures such as RNNs or hybrid models is preferred.
Benefits
- Series-A-stage ownership and influence on the company’s trajectory
- Work on large-scale model training systems and cutting-edge models
- Small, highly technical team with fast feedback loops
- Strong emphasis on engineering quality and research rigor
- Meaningful equity
- Competitive compensation
- Practical work arrangement and location are not stated
Tech Stack
Categories
About Featherless AI
Featherless AI builds a serverless inference platform that orchestrates GPUs and load balances models so teams can deploy and scale open‑source AI without managing infrastructure. Its public cloud serves tens of thousands of open‑weight models and supports fine‑tuning, targeting developers, ML engineers, and enterprises needing reliable, high‑throughput inference. Founded in 2023 and headquartered in San Francisco, the privately held, Series A company is backed by investors including AMD and Airbus Ventures.