1 day ago
Hsinchu, Taiwan or Taipei, TaiwanSenior / Staff+
Responsibilities
- Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
- Drive graph capture and deployment using PyTorch, ONNX, torch.compile, model rewrites, and graph-level transformations.
- Design and implement fusion kernels using DSL-based approaches such as Triton.
- Collaborate with compiler, performance, accuracy, and runtime teams on lowering strategies, kernel fusion, layout decisions, and integration.
- Profile and optimize LLM, VLM, and diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
- Own transformer optimizations involving KV cache management, decoding behavior, and long-context performance.
- Enable and optimize continuous batching and its effects on memory, scheduling, and tail latency.
- Architect distributed inference strategies such as sharding and parallelism across multi-core and multi-device systems.
- Create reusable approaches and tooling for scaling model optimization to new hardware architectures.
- Debug performance and stability issues, identify root causes, and drive production-ready solutions.
Requirements
- Expert expertise in PyTorch and inference-focused model optimization, with strong Python engineering skills.
- Hands-on experience with torch.compile, TorchDynamo, or related graph capture and compilation workflows.
- Deep understanding of transformer architectures, attention mechanisms, mixture-of-experts models, and performance trade-offs.
- Practical experience with KV cache behavior, serving-time optimization, and memory/performance trade-offs.
- Strong foundation in computer architecture, ML accelerators, and distributed systems.
- Ability to lead cross-functional technical efforts and influence design decisions.
- Bachelor’s degree with 4+ years, master’s degree with 3+ years, or PhD with 2+ years of relevant hardware, software, or systems engineering experience.
- Experience developing fusion kernels with Triton or similar DSLs and collaborating with ML compiler teams is preferred.
- Familiarity with LLM serving stacks and continuous batching systems is preferred.
- Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks is preferred.
- A PhD in a relevant field is preferred.
Tech Stack
Categories
About Qualcomm
Qualcomm is a public semiconductor company headquartered in San Diego, founded in 1985, that designs and sells wireless chipsets and platforms for mobile devices, automotive, IoT, and networking, notably the Snapdragon application processors and 5G modems. It also licenses a large portfolio of cellular patents to device makers, generating revenue alongside chip sales; its technology underpins many Android smartphones and emerging automotive and edge-compute systems, and it trades on NASDAQ as QCOM.
