1 day ago
Bengaluru, IndiaMid Level
Responsibilities
- Develop and optimize GPU compute kernels for OpenCL and Vulkan targeting AI and machine learning workloads.
- Design and extend MLIR dialects across frontend, graph, tensor, runtime, and low-level abstraction levels.
- Implement compiler passes and transformations including tiling, fusion, bufferization, vectorization, and lowering.
- Build GPU runtime infrastructure covering memory management, pipeline setup, command buffer orchestration, and resource scheduling.
- Develop code-generation pipelines that lower tensor IR through MLIR into OpenCL and Vulkan kernels.
- Profile compiled kernels using GPU counters and vendor-specific tools to identify bottlenecks and improve performance.
- Implement performance-critical schedules involving loop fusion, parallelism, and caching.
- Collaborate with framework teams to optimize model lowering for computer vision and LLM workloads.
- Develop compiler and runtime components in modern C/C++.
Requirements
- Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field with 4+ years of Systems Engineering or related experience; or a master's degree with 3+ years; or a PhD with 2+ years.
- Strong hands-on experience authoring MLIR dialects, writing compiler passes, and building end-to-end lowering pipelines.
- Expertise with MLIR frontend, graph-level, tensor IR, runtime, and low-level dialects, including TOSA, StableHLO, ONNX-MLIR, Linalg, Tensor, Vector, MemRef, SCF, GPU, and LLVM.
- Strong OpenCL programming experience, including kernel development, memory models, work-group and work-item optimization, and runtime management.
- Solid Vulkan compute programming experience, including descriptor management, compute pipelines, synchronization, and runtime internals.
- Strong understanding of GPU architecture, memory hierarchies, and asynchronous compute.
- Proficiency in C/C++ for system-level development and experience with GPU kernel profiling and bottleneck analysis.
- Strong machine learning fundamentals covering computer vision and LLM workloads.
- Preferred experience with IREE or MLIR-based frameworks such as TVM, XLA, or LLVM; IREE compiler and runtime architecture; open-source MLIR or IREE contributions; quantization; mixed-precision inference; model optimization; multi-target compilation; and tools such as ARM Streamline, Qualcomm Snapdragon Profiler, or Intel VTune.
About Qualcomm
Qualcomm is a public semiconductor company headquartered in San Diego, founded in 1985, that designs and sells wireless chipsets and platforms for mobile devices, automotive, IoT, and networking, notably the Snapdragon application processors and 5G modems. It also licenses a large portfolio of cellular patents to device makers, generating revenue alongside chip sales; its technology underpins many Android smartphones and emerging automotive and edge-compute systems, and it trades on NASDAQ as QCOM.
