Baseten

Forward Deployed Engineer (Training)

Baseten
Apply
4 hours ago

Base Salary

$200k - $400k/yr

Responsibilities

  • Own each customer account’s technical outcomes, including workload design, operation, and scaling on Baseten.
  • Translate vague customer objectives into specifications, success criteria, proofs of concept, and production solutions.
  • Design evaluations and benchmarks, optimize inference, improve models through post-training, and rework evaluations as needed.
  • Triage mission-critical failures, own or route fixes, and remain accountable through resolution.
  • Build evaluation and deployment tooling, automation, recipes, and reference implementations to improve future engagements.
  • Shape the product roadmap and ship fixes and features into Baseten’s codebase.
  • Manage multiple accounts, sequence work, coordinate stakeholders, and communicate status and risk.

Requirements

  • Minimum 1–2 years of software engineering experience shipping and maintaining code in large production systems, ideally across the stack.
  • Experience debugging complex production issues using logs, metrics, and traces to identify root causes in unfamiliar systems.
  • Ability to own ambiguous technical problems, triage issues, make decisions under uncertainty, and involve system owners when appropriate.
  • Interest in working directly with customers and influencing product direction beyond pure engineering responsibilities.
  • Ability to communicate complex technical topics with customer engineers, customer leadership, and internal stakeholders.
  • Curiosity about AI inference and training and motivation to develop expertise in AI infrastructure.
  • Willingness to respond to customers outside regular working hours and participate in an on-call rotation.
  • Depth in infrastructure domains such as storage or networking, including InfiniBand or RoCE, is relevant.
  • Experience operating distributed compute platforms such as Kubernetes, Slurm, or Ray, especially for GPU workloads, is relevant.
  • Understanding of LLM architectures and inference engines such as vLLM, TensorRT-LLM, or SGLang is relevant.
  • Ability to profile and optimize GPU workloads in training or serving is relevant.
  • Hands-on experience with post-training techniques such as SFT and RL, or deep learning experience with PyTorch or JAX, is relevant.
  • Operational experience with on-call work, incident response, and debugging distributed systems under pressure is relevant.

Benefits

  • Competitive compensation including meaningful equity.
  • 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible PTO and a company-wide Winter Break, with offices closed from Christmas Eve through New Year’s Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Exposure to a variety of ML startups and related learning and networking opportunities.

Categories

Forward DeployedML Engineering
Baseten

About Baseten

201-500 employees

Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.