
Technical Director, Large-Scale AI Model Inferencing
Samsung Semiconductor7 hours ago
San Jose, CA, USAStaff+
Base Salary
$219k - $351k/yr
Responsibilities
- Serve as the technical authority on how AI model architectures consume and move memory and translate that knowledge into memory-product requirements.
- Model memory footprint, bandwidth demand, and access patterns for dense Transformer, MoE, SSM, hybrid, diffusion, and multimodal models.
- Own production inference expertise and performance engineering across major inference stacks, including batching, caching, prefill/decode disaggregation, speculative decoding, and profiling.
- Define tiered memory requirements and policies across HBM, host DRAM, CXL-attached memory, and NVMe/SSD.
- Design expert-weight offloading and caching solutions for MoE model serving.
- Create requirements, reference architectures, analytical models, simulators, proofs of concept, and published benchmarks for AI memory products.
- Set multi-year technical strategy and build-vs-adopt-vs-contribute decisions across the inference and caching ecosystem.
- Lead architecture reviews, represent the company with customers and partners, contribute to standards discussions, and mentor senior engineers.
Requirements
- BS in Computer, Electrical, or Electronic Engineering or Computer Science with 20 years of relevant experience; an MS in one of these fields with 18 years of relevant experience is preferred.
- 12+ years of systems engineering experience and 4+ years of hands-on large-scale LLM inference or GPU systems performance experience.
- First-principles understanding of Transformer internals, KV-cache sizing, MQA, GQA, MLA, and activation-memory behavior.
- Working expertise with MoE routing, expert parallelism, load skew, and serving models larger than GPU capacity.
- Code-level experience with the memory-management internals of at least one major inference stack.
- Strong performance-engineering skills, including bandwidth-versus-compute analysis, NUMA and PCIe topology reasoning, RDMA basics, and GPU/CPU profiling.
- Production experience building systems software at the memory, storage, or I/O layer, such as caches, tiering, paging, or storage engines.
- Ability to develop analytical queueing, cache-hit-rate, and bandwidth models and simulators.
- Excellent written and verbal communication, including executive-level technical communication and customer-facing technical leadership.
- Preferred qualifications include SSM or hybrid serving experience, contributions to inference or caching open-source projects, CXL memory pooling, SSD/NVMe cache-tier engineering, and memory/storage product experience.
Benefits
- Base pay is listed separately; benefits include medical, dental, vision, and 401(k) coverage.
- 4+ weeks of paid time off per year, plus holidays and sick leave.
- Charitable giving match and community involvement opportunities.
- Fertility-care or adoption stipend, medical travel support, and virtual veterinary care.
- On-demand emotional wellness apps and free confidential therapy sessions.
- Onsite café and gym plus virtual fitness classes.
- Daily onsite presence is required at the San Jose headquarters, with a flexible work environment.
- The company provides accommodations throughout recruiting and supports an inclusive workplace.