d-Matrix

Principal LLM Inference Engineer

d-Matrix
Apply
3 months ago
Santa Clara, CA, USAStaff+
H1B sponsor

Base Salary

$195k - $285k/yr

Responsibilities

  • Identify and prototype emerging LLM inference use cases for heterogeneous hardware deployments.
  • Build proof-of-concept systems demonstrating D-Matrix capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels and operator-level optimizations to improve throughput and latency.
  • Drive quantization, sparsity, and batching strategies for the D-Matrix computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems, including tensor and pipeline parallelism, disaggregated prefill/decode, and KV-cache management.
  • Collaborate with hardware architects and provide firmware and compiler teams with inference workload insights.
  • Partner with product and business development to turn prototypes into customer-facing demonstrations.
  • Contribute to technical publications, whitepapers, and open-source projects.

Requirements

  • A bachelor’s degree in Computer Science, Electrical Engineering, or a related field with 10+ years of relevant engineering experience, or a master’s or PhD with 6+ years of relevant industry experience, or equivalent demonstrated experience.
  • Strong proficiency in Python and C/C++.
  • Hands-on experience optimizing LLM inference, including attention kernels, KV cache, batching strategies, and INT8, FP8, or INT4 quantization.
  • Contributor-level experience with at least one major inference framework such as vLLM, SGLang, TensorRT-LLM, or ONNX Runtime.
  • Familiarity with GPU kernel programming using CUDA or Triton and performance profiling tools.
  • Experience with heterogeneous compute deployments, custom silicon or ASIC-based inference, distributed inference, production inference serving at scale, or advanced techniques such as speculative decoding, mixture-of-experts routing, and long-context serving is preferred.
  • Open-source contributions to inference or ML systems projects and systems-level understanding of modern LLM training and inference are preferred.

Benefits

  • Competitive compensation, equity, and benefits are offered.
  • The role is based in Santa Clara, California.

Tech Stack

d-Matrix

About d-Matrix

201-500 employees

d-Matrix builds AI inference computing platforms for data centers, combining custom silicon with systems, networking, and software. Its flagship Corsair platform and JetStream fabric focus on low-latency, energy-efficient generative AI inference at scale. Founded in 2019 and headquartered in Santa Clara, California, the privately held company sells hardware with accompanying software to cloud providers and enterprises deploying large AI models.

Contact me