2 months ago
Base Salary
$200k - $300k/yr
Responsibilities
- Own the technical direction and roadmap for the ML infrastructure.
- Build and maintain the training and inference stack for fast experimentation and high-performance production serving.
- Optimize model serving across kernels, runtimes, batching, scheduling, and distributed inference.
- Design reliable multi-node, multi-GPU training and inference systems.
- Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
- Develop benchmarks to identify bottlenecks and guide infrastructure investments.
- Evaluate advances in training and inference and apply relevant improvements.
- Build tooling and abstractions that help ML engineers move experiments into production.
- Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
- Provide technical leadership through design reviews, mentorship, and hands-on implementation.
Requirements
- 5+ years of experience building production infrastructure, including significant machine learning systems experience.
- Experience leading complex technical projects from ambiguous problems through production deployment.
- Strong Python and systems-engineering skills.
- Understanding of the performance characteristics of modern GPU training or inference workloads.
- Experience with Kubernetes and distributed training or serving frameworks.
- Ability to reason across low-level model performance and higher-level platform architecture.
- Experience setting technical direction while personally implementing complex systems.
- Bonus: experience optimizing or implementing CUDA, Triton, or custom model-serving kernels.
- Bonus: meaningful contributions to vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
- Bonus: experience operating distributed inference or training across hundreds or thousands of GPUs.
- Bonus: experience building observability, scheduling, or capacity-management systems for GPU workloads.
- Bonus: experience at an early-stage or high-growth startup.
Benefits
- Fully in-person role at the San Francisco office.
- Unlimited PTO.
- Daily free lunch.
- Commuter reimbursement.
- Medical, dental, and vision insurance.
- Health and wellness budget of up to $150 per month.
- Flexible parental leave scheduling.
Tech Stack
Categories
About Reducto
Reducto is the complete agentic document platform for leading AI teams needing performance at enterprise scale. We provide a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows. We currently power the best AI teams ranging from startups to Fortune 10 companies across all industries: legal, finance, healthcare, and more. We are built for enterprise workloads with flexible deployment options from the cloud to fully air-gapped environments, SOC II and HIPAA compliance, and zero data retention.
