Inference

Senior Software Engineer - Model Performance

Inference
Apply
7 months ago

Base Salary

$220k - $320k/yr

Responsibilities

  • Implement and productionize inference optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving.
  • Debug and improve vLLM, SGLang, TensorRT-LLM, and underlying inference libraries.
  • Profile and optimize CUDA kernels and GPU utilization across the serving infrastructure.
  • Add support for new model architectures and validate their performance before production release.
  • Experiment with novel inference techniques and productionize successful approaches.
  • Build tooling and benchmarks to measure and track inference performance across the model-serving fleet.
  • Collaborate with applied ML engineers to ensure trained models can be served efficiently.

Requirements

  • At least 2 years of experience in ML systems, inference optimization, or GPU programming.
  • Strong proficiency in Python and familiarity with C++.
  • Hands-on experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar.
  • Deep understanding of GPU architecture and experience profiling GPU workloads.
  • Familiarity with quantization, speculative decoding, continuous batching, and KV cache management.
  • Experience with PyTorch and understanding of how models execute on hardware.
  • A track record of measurably improving system performance.
  • Preferred qualifications include CUDA programming, serving non-LLM models such as TTS, vision, or embeddings, distributed inference and multi-GPU serving, contributions to open-source inference frameworks, and Docker and Kubernetes experience.

Benefits

  • Equity and comprehensive benefits are offered.
  • The role is primarily in-person in downtown San Francisco, with hybrid work available for Bay Area candidates; most employees work in the office four days per week.
Inference

About Inference

11-50 employees

Inference.net helps teams ship AI that’s faster, smarter, and dramatically more cost-efficient. We deliver lower latency and higher-quality models at a fraction of the cost, with full OpenAI compatibility and no vendor lock-in. Companies use Inference.net to power real-time AI features, automate workflows, and scale mission-critical systems without blowing up their margins. Case studies: • The fastest growing AEO optimization software: trained a custom model that improved ranking accuracy for Fortune 500 clients, increasing conversion while lowering inference spend and stabilizing latency across high-volume workloads. • Fastest-growing nutrition tracking app: scaled to 10M+ users while cutting inference costs, processing millions of daily food images with custom vision models outperforming frontier VLMs. • Decentralized data network (search.video): processed billions of monthly video frames across 3M+ nodes using a specialized ClipTagger that boosted relevance and throughput. • Digital bank (120M+ customers): delivered 99.99%+ uptime, faster and more accurate compliance/servicing models at a fraction of API cost. Build AI products that scale, without sacrificing performance or profitability. You can find us also on X: https://x.com/inference_net