about 3 hours ago
New York, NY, USAMid Level / Senior
Base Salary
$150k - $350k/yr
Responsibilities
- Maintain and enhance the inference stack on-device and in the cloud.
- Integrate new model architectures and multimodal models.
- Optimize latency, throughput, memory use, and reliability across various runtimes.
- Develop runtime capabilities for model loading, batching, scheduling, caching, and distributed execution.
- Benchmark and diagnose performance issues across the inference stack.
- Contribute to open-source projects like llama.cpp and MLX.
Requirements
- Significant experience in building production ML systems or inference runtimes.
- Strong programming skills in Python and C++.
- Deep understanding of transformer architectures and model inference mechanics.
- Experience profiling CPU or GPU workloads and understanding compute and memory management.
- Familiarity with PyTorch and inference systems such as llama.cpp, MLX, or TensorRT-LLM.
- Strong debugging skills across model code and runtime internals.
Benefits
- Competitive salary and equity grants.
- Comprehensive medical, vision, and dental healthcare plans.
- Catered team lunches and expensed dinners in the office.
- Flexible PTO and work-from-home options.
- Sun-drenched office located in SoHo, NYC.
- Access to powerful hardware for development.
