7 days ago
Base Salary
$200k - $320k/yr
Responsibilities
- Own inference performance across latency, throughput, and cost per unit of work.
- Optimize kernels, batching, quantization, compilation, serving, and hardware utilization to reduce compute costs.
- Profile inference pipelines with Nsight, PyTorch Profiler, or equivalent tools and ship performance fixes.
- Partner with research and product teams to make accurate models affordable to run at scale.
- Build serving and evaluation paths that expose inference costs during experimentation.
- Track cost as a first-class metric alongside model quality.
Requirements
- Strong ML or systems engineering background with real inference optimization experience in areas such as serving, compilers, CUDA, kernels, or quantization.
- Proficiency in Python and C++ or Rust for performance-critical paths.
- Familiarity with PyTorch or JAX and profiling tools.
- Ability to reason about dollars, tokens or frames per second, memory hierarchy, data movement, and low-precision compute.
- Experience working in a small research team and shipping under cost pressure.
- Production experience with CUDA, kernels, compilers such as TVM, MLIR, or TensorRT, quantization, GPU or accelerator cost optimization, or serving stacks for video or large models is a plus.
Benefits
- Medical, dental, and vision packages with generous premium coverage.
- $500 per month credit for waiving medical benefits.
- $2k per month housing subsidy for employees living within walking distance of the office.
- Relocation support for moves to San Francisco or Shenzhen.
- Wellness benefits covering fitness and mental health.
- Daily lunch and dinner at the office.
- Unlimited compute budget subject to ROI justification.
- Unlimited Codex and Claude credits.
- Travel benefits.
- Fully in-person work arrangement in San Francisco's Financial District or Shenzhen's Nanshan district.