5 months ago
Base Salary
$180k - $280k/yr
Responsibilities
- Develop, enhance, and maintain next-generation AI deployment software.
- Build and scale deployment infrastructure for the AI compute engine within tight development windows.
- Work across the full-stack toolchain and optimize hardware-software co-design trade-offs.
- Collaborate with system software, machine learning, compiler, and hardware experts.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Math, Physics, or a related field is required, with 12+ years of industry software development experience.
- Master’s degree or PhD in Computer Science, Electrical Engineering, or a related field is preferred.
- Strong understanding of system software, data structures, computer architecture, and machine learning fundamentals is required.
- Proficiency in C, C++, and Python development in a Linux environment using standard development tools is required.
- Experience designing and implementing distributed, high-performance software is required.
- Preferred experience includes TensorRT-LLM, vLLM, SGLang, PyTorch, TensorFlow, ONNX Runtime, TensorRT, NCCL, OpenMPI, Kubernetes, and Ray.
- Preferred experience includes deploying LLM, VLM, and NLP workloads on distributed systems and using MLOps tools from definition through deployment.
- Experience with software testing fundamentals, startup or small-team environments, or cloud providers and AI compute/subsystem companies is preferred.
Benefits
- Hybrid work arrangement with onsite work at the Santa Clara, California headquarters three days per week.
Tech Stack
Categories
About d-Matrix
d-Matrix builds AI inference computing platforms for data centers, combining custom silicon with systems, networking, and software. Its flagship Corsair platform and JetStream fabric focus on low-latency, energy-efficient generative AI inference at scale. Founded in 2019 and headquartered in Santa Clara, California, the privately held company sells hardware with accompanying software to cloud providers and enterprises deploying large AI models.
