19 days ago
Base Salary
$230k - $300k/yr
Responsibilities
- Lead the technical direction and architecture of the model serving platform.
- Build execution runtimes, batching and scheduling systems, and distributed inference components.
- Develop high-performance C++ and CUDA/HIP modules, custom GPU kernels, and memory-optimized runtimes.
- Collaborate with ML researchers to productionize multimodal models for low-latency, scalable inference.
- Build Python APIs and services for downstream applications.
- Mentor engineers through code reviews, design discussions, and hands-on technical guidance.
- Drive performance profiling, benchmarking, and observability across the inference stack.
- Maintain reliability and troubleshoot complex issues across GPU, runtime, and service layers.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- 5+ years of experience designing and building scalable, reliable backend systems or distributed infrastructure.
- Strong understanding of LLM inference mechanics, including prefill, decode, batching, and KV cache.
- Experience with Kubernetes, Ray, and containerization.
- Strong proficiency in C++ and Python.
- Strong system-level debugging, profiling, and performance optimization skills.
- Ability to collaborate with ML researchers and translate model or runtime requirements into production systems.
- Ability to lead technical discussions, mentor engineers, and drive engineering quality.
- Preferred experience includes ML systems engineering, distributed GPU scheduling, large-scale ML/MLOps infrastructure, CUDA or ROCm, GPU profiling tools, multimodal model architectures, efficient inference techniques, and open-source ML or HPC infrastructure contributions.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
- Work from the office.
