5 months ago
Base Salary
$200k - $275k/yr
Responsibilities
- Build in-house tooling to support customer post-training of custom models.
- Support training a broad range of model architectures with varied techniques efficiently and at scale.
- Work across systems-level infrastructure, including Kubernetes, cgroups, storage systems, and networking topologies.
- Develop and optimize PyTorch distributed tensor computations and GPU kernels.
- Investigate deep technical topics and collaborate with researchers to solve complex post-training problems.
Requirements
- Deep understanding of modern machine learning techniques and tools for training transformers.
- Advanced experience with a tensor or array computation library such as PyTorch, TensorFlow, or Jax.
- Detailed understanding of data, sharded data, tensor, pipeline, and context parallelism strategies for transformer training.
- Experience profiling and improving the performance of distributed GPU programs in PyTorch or a similar library.
- Ability to perform roofline analysis on a transformer training setup.
- Familiarity with HPC and distributed computing platforms such as Slurm, Ray, Kubernetes, and Dask.
- Familiarity with cluster networking technologies including Infiniband, RoCE, and GPUDirect.
- Solid operating systems fundamentals covering processes, files, kernel drivers, containerization, and networking protocols.
- Willingness to work through ambiguous problems with researchers, derive specifications, and execute.
- Creativity and willingness to question approaches, assumptions, and tooling choices.
Benefits
- Competitive compensation including meaningful equity.
- Medical, dental, and vision insurance fully covered for employees and dependents.
- Flexible PTO and a company-wide Winter Break from Christmas Eve through New Year's Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of machine learning startups and networking opportunities.
Tech Stack
Categories
AI ResearchML Engineering
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
