2 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Maintain and advance the inference stack for on-device and cloud environments.
- Bring up new model architectures and multimodal models.
- Improve latency, throughput, memory usage, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes.
- Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution.
- Benchmark and diagnose correctness and performance issues across the inference stack.
- Contribute upstream improvements to open-source projects such as llama.cpp and MLX.
Requirements
- Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure.
- Strong programming ability in Python and C++.
- Deep understanding of transformer architectures and model inference mechanics.
- Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement.
- Experience with PyTorch and inference systems including llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM.
- Strong debugging ability across model code, runtime internals, operating systems, and CPU or GPU execution.
- Open-source contributions to inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM are preferred.
Benefits
- Competitive salary and equity grants; compensation figures are not stated.
- Medical, vision, and dental healthcare plans.
- Catered team lunches and expensed dinners in the office.
- Flexible paid time off.
- Flexible work-from-home arrangement.
- Office located in SoHo in New York City.
- Access to powerful hardware.
- The posting does not state a contract or internship duration.
Categories
About LM Studio
Run LLMs on your own hardware. Download the community version: https://lmstudio.ai/download Enterprise: https://lmstudio.ai/enterprise
