10 months ago
Base Salary
$190k - $250k/yr
Responsibilities
- Build the model-serving platform, including API, control plane, billing, monitoring, and distributed inference features.
- Integrate new multimodal models into production workflows in collaboration with ML researchers.
- Write reliable, maintainable, well-tested, and documented code.
- Provide operational support to keep production services performant, available, and reliable.
- Troubleshoot complex issues across runtime, service, and GPU layers with other engineers.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- At least 3 years of software engineering experience focused on infrastructure or machine learning systems.
- Strong proficiency in C++, Python, Go, or Rust.
- Experience with Kubernetes and containerization.
- Experience building large-scale ML or MLOps infrastructure.
- Strong collaboration and communication skills across engineering and ML teams.
- Nice to have experience with ML systems engineering and open-source inference engines such as vLLM, Sglang, or TRT-LLM.
- Nice to have proficiency in CUDA or ROCm and experience with GPU profiling tools.
- Nice to have contributions to open-source ML or HPC infrastructure.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
- Work from the office.
