1 year ago
Base Salary
$165k - $330k/yr
Responsibilities
- Design and architect scalable infrastructure systems for the ML training platform, including scheduling, storage, and networking.
- Partner with developers and research engineers to translate complex training requirements into technical solutions.
- Design and architect a global training scheduler.
- Design and architect reinforcement learning systems and continuous learning pipelines.
- Drive long-term improvements to system reliability and development velocity.
- Partner with SRE and Capacity teams to advance training infrastructure.
- Make critical architectural decisions that balance performance and reliability.
- Lead technical discussions and mentor junior engineers on infrastructure best practices.
- Contribute to long-term technical strategy and the infrastructure roadmap.
Requirements
- Bachelor’s degree or higher in Computer Science or a related field.
- Proficiency in Go; Python experience is a plus.
- Deep expertise with Kubernetes in production environments.
- Extensive experience with major cloud providers such as AWS and GCP; experience with Crusoe, DigitalOcean, or Nebius is a plus.
- Advanced understanding of distributed systems concepts and performance tuning.
- Proven experience designing observability systems.
- Experience with ML/AI workloads and MLOps platforms is highly valued.
- Experience with distributed storage systems is a nice-to-have.
- Experience with workload orchestration platforms such as Temporal or Airflow is a nice-to-have.
- Familiarity or experience with NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainier, FSDP, and DeepSpeed is a nice-to-have.
- Experience developing AI products, tooling, or agents is a nice-to-have.
Benefits
- 100% coverage of medical, dental, and vision insurance for employees and dependents
- Flexible PTO and company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups and networking opportunities
Tech Stack
Categories
About Baseten
Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.
