Baseten

Software Engineer- Inference Platform

Baseten
Apply
4 hours ago
Toronto, Canada +4 moreMid Level
H1B sponsor

Base Salary

$180k - $360k/yr

Responsibilities

  • Build infrastructure and orchestration systems for large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
  • Design, build, and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
  • Implement API versioning, validation, usage metering, quotas, and authentication.
  • Build observability, benchmarks, testing and release practices, and operational processes for speed, reliability, and quality.
  • Debug and harden production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Partner with inference performance engineers to make optimizations broadly available and easy for customers to configure.
  • Own projects end to end from architecture through deployment, monitoring, and iteration based on customer feedback.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • At least 3 years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs with reliability, latency, and scale requirements.
  • Proven experience owning low-latency, reliable backend services involving rate limiting, authentication, quotas, metering, and migrations.
  • Experience with profiling, tracing, capacity planning, and SLO management.
  • Ability to debug performance and reliability issues across application, runtime, and infrastructure layers.
  • Strong developer-experience orientation and interest in inference engineering; prior ML or LLM experience is not required.
  • Excellent written communication and collaboration skills, including design documentation and cross-functional work.
  • Preferred experience with LLM inference engines or frameworks such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo.
  • Preferred deep Kubernetes experience, including operators and custom resources, plus familiarity with service meshes or API gateways.
  • Preferred experience with distributed scheduling, autoscaling, service orchestration, production GPU workloads, developer-facing infrastructure or APIs, open-source infrastructure or ML systems, observability tooling, CI/CD systems, or release automation.

Benefits

  • Competitive compensation including meaningful equity.
  • For U.S. employees and dependents, 100% coverage of medical, dental, and vision insurance.
  • Flexible PTO and a company-wide Winter Break from Christmas Eve through New Year's Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • For U.S. employees, a company-facilitated 401(k).
  • Exposure to a variety of ML startups and associated learning and networking opportunities.
Baseten

About Baseten

201-500 employees

Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.

Contact me