
Software Engineer, Production Inference (Distributed Inference)
Thinking Machines Lab4 hours ago
Base Salary
$350k - $500k/yr
Responsibilities
- Design, build, and operate distributed infrastructure for large-scale model serving, including request routing, load balancing, batching, and multi-node coordination.
- Optimize production inference latency and throughput through KV cache management, continuous batching, speculative decoding, and quantization.
- Build and maintain high-concurrency serving systems with strong uptime, low tail latency, and deep observability.
- Benchmark, tune, and extend inference engines for new model architectures moving from research into production.
- Partner with research and infrastructure teams to translate emerging model designs into production-ready serving systems.
- Build tracing and debugging tooling across the serving stack, from orchestration through GPU kernels.
- Participate in an on-call rotation supporting production inference systems.
Requirements
- At least 3 years of experience building and operating distributed systems in production.
- Strong systems programming skills in Python, C++, Rust, or similar languages.
- Experience with production infrastructure at scale, including reliability, observability, and performance under real-world load.
- Solid understanding of networking, concurrency, and distributed-systems fundamentals.
- Preferred: experience with LLM inference engines such as vLLM, SGLang, or TensorRT-LLM.
- Preferred: familiarity with GPU programming using CUDA or low-level performance optimization.
- Preferred: experience with model parallelism, tensor parallelism, pipeline parallelism, or other distributed inference techniques.
- Preferred: experience operating large-scale production systems with strict latency and uptime requirements.
- Preferred: experience with Kubernetes or similar orchestration systems for GPU workloads.
- Preferred: contributions to open-source ML systems or inference infrastructure projects.
Benefits
- The role is based in San Francisco, California.
- The company offers health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.