Samsung Semiconductor

Technical Director, Large-Scale AI Model Inferencing

Samsung Semiconductor
Apply
7 hours ago
San Jose, CA, USAStaff+

Base Salary

$219k - $351k/yr

Responsibilities

  • Serve as the technical authority on how AI model architectures consume and move memory and translate that knowledge into memory-product requirements.
  • Model memory footprint, bandwidth demand, and access patterns for dense Transformer, MoE, SSM, hybrid, diffusion, and multimodal models.
  • Own production inference expertise and performance engineering across major inference stacks, including batching, caching, prefill/decode disaggregation, speculative decoding, and profiling.
  • Define tiered memory requirements and policies across HBM, host DRAM, CXL-attached memory, and NVMe/SSD.
  • Design expert-weight offloading and caching solutions for MoE model serving.
  • Create requirements, reference architectures, analytical models, simulators, proofs of concept, and published benchmarks for AI memory products.
  • Set multi-year technical strategy and build-vs-adopt-vs-contribute decisions across the inference and caching ecosystem.
  • Lead architecture reviews, represent the company with customers and partners, contribute to standards discussions, and mentor senior engineers.

Requirements

  • BS in Computer, Electrical, or Electronic Engineering or Computer Science with 20 years of relevant experience; an MS in one of these fields with 18 years of relevant experience is preferred.
  • 12+ years of systems engineering experience and 4+ years of hands-on large-scale LLM inference or GPU systems performance experience.
  • First-principles understanding of Transformer internals, KV-cache sizing, MQA, GQA, MLA, and activation-memory behavior.
  • Working expertise with MoE routing, expert parallelism, load skew, and serving models larger than GPU capacity.
  • Code-level experience with the memory-management internals of at least one major inference stack.
  • Strong performance-engineering skills, including bandwidth-versus-compute analysis, NUMA and PCIe topology reasoning, RDMA basics, and GPU/CPU profiling.
  • Production experience building systems software at the memory, storage, or I/O layer, such as caches, tiering, paging, or storage engines.
  • Ability to develop analytical queueing, cache-hit-rate, and bandwidth models and simulators.
  • Excellent written and verbal communication, including executive-level technical communication and customer-facing technical leadership.
  • Preferred qualifications include SSM or hybrid serving experience, contributions to inference or caching open-source projects, CXL memory pooling, SSD/NVMe cache-tier engineering, and memory/storage product experience.

Benefits

  • Base pay is listed separately; benefits include medical, dental, vision, and 401(k) coverage.
  • 4+ weeks of paid time off per year, plus holidays and sick leave.
  • Charitable giving match and community involvement opportunities.
  • Fertility-care or adoption stipend, medical travel support, and virtual veterinary care.
  • On-demand emotional wellness apps and free confidential therapy sessions.
  • Onsite café and gym plus virtual fitness classes.
  • Daily onsite presence is required at the San Jose headquarters, with a flexible work environment.
  • The company provides accommodations throughout recruiting and supports an inclusive workplace.

Categories

Samsung Semiconductor

About Samsung Semiconductor

10,000+ employees
Contact me