1 day ago
Base Salary
$209k - $283k/yr
Responsibilities
- Define architecture, interfaces, and roadmap for distributed AI inference runtime capabilities.
- Enable new model architectures through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.
- Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.
- Evaluate inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.
- Partner with cloud, framework, compiler, hardware, research, compute, and product teams.
- Lead technical reviews, mentor engineers, and establish performance-engineering practices.
Requirements
- At least 5 years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
- Deep understanding of AI inference, including model execution, attention, mixture-of-experts, batching, prioritization, and KV-cache behavior.
- Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, memory resources, and data movement.
- Ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
- Preferred experience with inference schedulers, cache managers, batching systems, and disaggregated or distributed execution paths.
- Preferred experience optimizing kernels with accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.
- Preferred familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.
- Preferred contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.
Benefits
- Competitive salary and total reward package, with details shared during recruitment.
- Hybrid working with team-determined patterns and role-specific flexibility details provided during application.
- Accommodations and adjustments are available during the recruitment process.
- Collaborative, diverse environment focused on foundational production AI infrastructure.
Categories
About Arm
Arm’s foundational technology is defining the future of computing. A future built by the greatest technology ecosystem in the world. A future built on Arm. Arm is everywhere technology matters. Technology matters everywhere. Together, we’ll power every technology revolution moving forward, including cloud computing, automotive and autonomous systems, IoT, the metaverse, and beyond. Changing the world. Again. On Arm.
