Baseten

Software Engineer- Inference Performance

Baseten
Apply
3 hours ago
Toronto, Canada +4 moreSenior
H1B sponsor

Base Salary

$180k - $360k/yr

Responsibilities

  • Implement and productionize inference techniques in runtime internals, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, scheduling, and routing.
  • Profile and optimize inference end to end, from kernel launch overhead and memory layout through request scheduling, batching, and cache-aware routing.
  • Improve tokens per GPU-hour, utilization, latency, throughput, and serving cost while communicating performance tradeoffs.
  • Bring up and tune new model architectures on new hardware.
  • Build benchmarking frameworks across model architectures, batch sizes, sequence lengths, and hardware configurations.
  • Contribute to open-source inference engines and collaborate with model, infrastructure, and customer-facing teams.

Requirements

  • Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
  • Experience with a general-purpose programming language such as Python or C++.
  • Familiarity with LLM optimization techniques including quantization, speculative decoding, and continuous batching.
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.
  • Preferred qualifications include inference-engine contributions, distributed serving experience, GPU-kernel optimization, quantization or speculative decoding in production, and experience developing and deploying AI/ML inference solutions.

Benefits

  • Competitive compensation with meaningful equity.
  • In the U.S., 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible PTO and a company-wide Winter Break from Christmas Eve through New Year’s Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • In the U.S., a company-facilitated 401(k).
  • Exposure to a variety of ML startups and associated learning and networking opportunities.
Baseten

About Baseten

201-500 employees

Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.

Contact me