1 day ago
Bengaluru, IndiaSenior / Staff+
H1B Sponsor
Responsibilities
- Identify and characterize critical ML models and use cases for CPU execution, including compute intensity, memory behavior, parallelism, and dataflow.
- Generate execution traces with QEMU or equivalent simulators and develop tooling to capture instruction behavior, performance counters, and bottlenecks.
- Analyze CPU pipeline, memory hierarchy, and instruction-utilization bottlenecks and optimize hotspots through kernel tuning, algorithmic improvements, and memory optimization.
- Collaborate with CPU architecture and design teams to recommend architectural enhancements based on real workload data.
- Design and implement optimized QMX ML kernels and libraries for GEMM, convolution, attention, activation functions, and related workloads.
- Integrate optimized kernels with open-source ML frameworks and inference stacks.
- Optimize CPU-centric ML benchmarks, establish performance baselines, track improvements across hardware generations, and perform competitive analysis.
Requirements
- Bachelor’s degree in Engineering, Information Systems, Computer Science, or a related field plus 8+ years of software engineering or related experience; or a master’s degree plus 7+ years; or a PhD plus 6+ years.
- At least 4 years of experience with programming languages such as C, C++, Java, or Python.
- Strong background in computer architecture, systems programming, and machine learning fundamentals.
- Mandatory proficiency in C/C++.
- Experience with performance profiling, benchmarking, and optimization.
- Preferred experience with QEMU or equivalent simulators and ML kernel development for GEMM, convolution, and attention.
- Knowledge of CPU pipelines, caching, SIMD/vector extensions such as NEON, SVE, and QMX, plus familiarity with ML frameworks and inference stacks.
- Experience with intrinsics, assembly, memory tuning, and cache tuning.
Benefits
- Opportunity to work on next-generation Qualcomm CPU architectures and directly influence hardware design through workload insights.
- Collaboration with architecture, systems, and AI teams with visibility across product and research roadmaps.
- Role location is Bangalore or a relevant location.
- Qualcomm provides reasonable workplace and hiring-process accommodations for individuals with disabilities.
About Qualcomm
Inspired by 250 years of American ingenuity, we build technologies that expand what’s possible. From wireless to AI at scale and 6G, our innovations power progress worldwide.
