d-Matrix

Principal LLM Inference Engineer

d-Matrix
Apply
2 months ago
Santa Clara, CA, USAStaff+
H1B Sponsor

Base Salary

$195k - $285k/yr

Responsibilities

  • Identify and prototype emerging LLM inference use cases for heterogeneous hardware deployments.
  • Build proof-of-concept systems demonstrating D-Matrix capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels and operator-level optimizations to improve throughput and latency.
  • Drive quantization, sparsity, and batching strategies for the D-Matrix computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems, including tensor and pipeline parallelism, disaggregated prefill/decode, and KV-cache management.
  • Collaborate with hardware architects and provide firmware and compiler teams with inference workload insights.
  • Partner with product and business development to turn prototypes into customer-facing demonstrations.
  • Contribute to technical publications, whitepapers, and open-source projects.

Requirements

  • A bachelor’s degree in Computer Science, Electrical Engineering, or a related field with 10+ years of relevant engineering experience, or a master’s or PhD with 6+ years of relevant industry experience, or equivalent demonstrated experience.
  • Strong proficiency in Python and C/C++.
  • Hands-on experience optimizing LLM inference, including attention kernels, KV cache, batching strategies, and INT8, FP8, or INT4 quantization.
  • Contributor-level experience with at least one major inference framework such as vLLM, SGLang, TensorRT-LLM, or ONNX Runtime.
  • Familiarity with GPU kernel programming using CUDA or Triton and performance profiling tools.
  • Experience with heterogeneous compute deployments, custom silicon or ASIC-based inference, distributed inference, production inference serving at scale, or advanced techniques such as speculative decoding, mixture-of-experts routing, and long-context serving is preferred.
  • Open-source contributions to inference or ML systems projects and systems-level understanding of modern LLM training and inference are preferred.

Benefits

  • Competitive compensation, equity, and benefits are offered.
  • The role is based in Santa Clara, California.

Tech Stack

d-Matrix

About d-Matrix

201-500 employees

d-Matrix is an AI infrastructure company building the next generation of inference computing for the era of generative and agentic AI. Founded in 2019, d-Matrix is rethinking AI inference from the ground up with a full-stack approach spanning silicon, systems, networking, and software. Its flagship products, including the Corsair™ inference platform and JetStream™ inference fabric, are purpose-built to deliver high-performance, low-latency, and energy-efficient AI inference at datacenter scale. The company has raised nearly $500 million from a global syndicate of leading venture capital firms, sovereign wealth funds, and strategic investors, including Playground Global, Bullhound Capital, M12 (Microsoft’s Venture Fund), SK hynix, Temasek, Qatar Investment Authority, and Singapore’s EDBI. The company’s most recent financing valued d-Matrix at approximately $2 billion. As AI shifts from training to inference, the demands on AI infrastructure are changing. d-Matrix is building the infrastructure required to power the next generation of real-time AI at scale.