Baseten

Software Engineer, Model Performance Systems

Baseten
Apply
7 months ago
San Francisco, CA, USA or New York, NY, USAEntry Level
H1B Sponsor

Base Salary

$160k - $200k/yr

Responsibilities

  • Run and automate LLM quality benchmarks including GSM8K and MMLU, along with custom workload performance suites.
  • Create automated acceptance tests for GPU clusters across x86 and ARM systems, measuring GPU memory bandwidth, network throughput, and multi-node performance.
  • Develop and maintain internal GPU-enabled development environments for model experimentation.
  • Build and contribute to tools such as InferenceMAX and genai-bench for model evaluation and optimization.
  • Use PyTorch Profiler and NVIDIA Nsight Systems to profile performance, identify bottlenecks, and debug NVIDIA compute and networking systems.
  • Develop real-time dashboards and alerts for system health, model startup times, and runtime performance.
  • Automate performance testing through CI/CD pipelines to detect model setup regressions before production.
  • Build optimization tools that identify the best latency, cost, and quality configurations for models and workloads.

Requirements

  • Interest in GPU memory systems, InfiniBand, cluster networking, and how data moves across hardware.
  • An automation mindset and interest in stress-testing and fuzz testing systems.
  • Curiosity about the mathematics of Transformers, FLOPs, and memory requirements.
  • Interest in quantization, speculative decoding, disaggregated serving, and kernel-level optimization.
  • Familiarity with Python and eagerness to learn the NVIDIA software stack.
  • C++ familiarity is preferred.
  • The role is open to early-career and fresher candidates, with trajectory, curiosity, and technical depth emphasized over years of experience.

Benefits

  • Meaningful equity and competitive compensation are offered, without a stated amount.
  • Medical, dental, and vision insurance are fully covered for employees and dependents.
  • Flexible PTO includes a company-wide Winter Break from Christmas Eve through New Year's Day.
  • Paid parental leave is provided.
  • Fertility and family-building support is available through Carrot.
  • Company-facilitated 401(k) is available.
  • The role offers exposure to a variety of ML startups and learning and networking opportunities.

Tech Stack

Baseten

About Baseten

201-500 employees

Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.