d-Matrix

Principal Architect, Performance Analysis and Modeling

d-Matrix
Apply
2 months ago
Remote, United States or Santa Clara, CA, USAStaff+
H1B sponsor

Base Salary

$175k - $285k/yr

Responsibilities

  • Analyze emerging ML workloads, multimodal LLMs, chain-of-thought reasoning models, and video/audio generation for performance-relevant properties.
  • Build and maintain analytical performance models for current and future d-Matrix hardware generations.
  • Develop and extend architecture simulators for proposed hardware/software features.
  • Collaborate with hardware design, compiler, inference server, kernel, and product teams to validate modeling assumptions and identify downstream implications.
  • Track AI/ML architecture and algorithm research and incorporate relevant findings into modeling work.
  • Propose hardware/software feature improvements based on modeling and workload-analysis results.
  • Document modeling methodologies and findings for reuse across the architecture team.

Requirements

  • BSEE with 6+ years of industry experience or MSEE with 4+ years of industry experience.
  • Working knowledge of computer architecture, hardware/software co-design, performance modeling, and ML fundamentals, particularly DNNs.
  • Programming fluency in C, C++, or Python.
  • Experience building or working with analytical performance models or architecture simulators.
  • Experience optimizing AI/ML workloads on accelerator technologies is preferred.
  • Research or investigation experience in AI/ML architecture or microarchitecture is preferred.
  • Self-motivated and collaborative, with comfort working across hardware and software teams.

Benefits

  • Hybrid onsite schedule in Santa Clara, California, three days per week.
  • Remote work within the United States may be considered.
  • Equal opportunity and affirmative action workplace.

Tech Stack

d-Matrix

About d-Matrix

201-500 employees

d-Matrix builds AI inference computing platforms for data centers, combining custom silicon with systems, networking, and software. Its flagship Corsair platform and JetStream fabric focus on low-latency, energy-efficient generative AI inference at scale. Founded in 2019 and headquartered in Santa Clara, California, the privately held company sells hardware with accompanying software to cloud providers and enterprises deploying large AI models.

Contact me