12 months ago
Base Salary
$165k - $330k/yr
Responsibilities
- Design and architect scalable infrastructure systems for the ML training platform, including scheduling, storage, and networking.
- Partner with developers and research engineers to translate complex training requirements into technical solutions.
- Design and architect a global training scheduler.
- Design and architect reinforcement learning systems and continuous learning pipelines.
- Drive long-term improvements to system reliability and development velocity.
- Partner with SRE and Capacity teams to advance training infrastructure.
- Make critical architectural decisions that balance performance and reliability.
- Lead technical discussions and mentor junior engineers on infrastructure best practices.
- Contribute to long-term technical strategy and the infrastructure roadmap.
Requirements
- Bachelor’s degree or higher in Computer Science or a related field.
- Proficiency in Go; Python experience is a plus.
- Deep expertise with Kubernetes in production environments.
- Extensive experience with major cloud providers such as AWS and GCP; experience with Crusoe, DigitalOcean, or Nebius is a plus.
- Advanced understanding of distributed systems concepts and performance tuning.
- Proven experience designing observability systems.
- Experience with ML/AI workloads and MLOps platforms is highly valued.
- Experience with distributed storage systems is a nice-to-have.
- Experience with workload orchestration platforms such as Temporal or Airflow is a nice-to-have.
- Familiarity or experience with NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainier, FSDP, and DeepSpeed is a nice-to-have.
- Experience developing AI products, tooling, or agents is a nice-to-have.
Benefits
- 100% coverage of medical, dental, and vision insurance for employees and dependents
- Flexible PTO and company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups and networking opportunities
Tech Stack
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
