7 months ago
Base Salary
$160k - $200k/yr
Responsibilities
- Run and automate LLM quality benchmarks including GSM8K and MMLU, along with custom workload performance suites.
- Create automated acceptance tests for GPU clusters across x86 and ARM systems, measuring GPU memory bandwidth, network throughput, and multi-node performance.
- Develop and maintain internal GPU-enabled development environments for model experimentation.
- Build and contribute to tools such as InferenceMAX and genai-bench for model evaluation and optimization.
- Use PyTorch Profiler and NVIDIA Nsight Systems to profile performance, identify bottlenecks, and debug NVIDIA compute and networking systems.
- Develop real-time dashboards and alerts for system health, model startup times, and runtime performance.
- Automate performance testing through CI/CD pipelines to detect model setup regressions before production.
- Build optimization tools that identify the best latency, cost, and quality configurations for models and workloads.
Requirements
- Interest in GPU memory systems, InfiniBand, cluster networking, and how data moves across hardware.
- An automation mindset and interest in stress-testing and fuzz testing systems.
- Curiosity about the mathematics of Transformers, FLOPs, and memory requirements.
- Interest in quantization, speculative decoding, disaggregated serving, and kernel-level optimization.
- Familiarity with Python and eagerness to learn the NVIDIA software stack.
- C++ familiarity is preferred.
- The role is open to early-career and fresher candidates, with trajectory, curiosity, and technical depth emphasized over years of experience.
Benefits
- Meaningful equity and competitive compensation are offered, without a stated amount.
- Medical, dental, and vision insurance are fully covered for employees and dependents.
- Flexible PTO includes a company-wide Winter Break from Christmas Eve through New Year's Day.
- Paid parental leave is provided.
- Fertility and family-building support is available through Carrot.
- Company-facilitated 401(k) is available.
- The role offers exposure to a variety of ML startups and learning and networking opportunities.
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
