Arm

Principal Software Engineer, AI Inference Runtime

Arm
Apply
1 day ago

Base Salary

$263k - $355k/yr

Responsibilities

  • Define the architecture, interfaces, and roadmap for AI inference runtime capabilities.
  • Enable new model architectures through operator support, production validation, scheduling, batching, model execution, memory management, and KV-cache optimization.
  • Profile bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.
  • Evaluate inference techniques and build benchmarking, regression, validation, and safe-rollout systems.
  • Partner with cloud, framework, compiler, hardware, research, infrastructure, and product teams.
  • Lead technical reviews, mentor engineers, and establish performance-engineering practices.

Requirements

  • 8+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
  • Deep understanding of AI inference, including model execution, Attention, Mixture of Experts, batching, prioritisation, and KV-cache behavior.
  • Strong programming skills in C++, Rust, Python, or a comparable language.
  • Knowledge of concurrency, parallel programming, memory resources, and data movement.
  • Ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
  • Experience with inference schedulers, cache managers, batching systems, or disaggregated and distributed execution paths is preferred.
  • Experience with accelerator programming tools, assembly, intrinsics, attention, matrix multiplication, operator fusion, or low-precision execution is preferred.
  • Familiarity with model parallelism, collective communication, high-performance networking, compilers, graph optimization, or open-source ML runtimes, frameworks, compilers, and kernel libraries is preferred.

Benefits

  • Salary range of $262,700-$355,400 per year, with total rewards shared during recruitment.
  • Hybrid working arrangements determined by the team, with role-specific details provided during application.
  • Accommodation and adjustment support is available during the recruitment process.
  • Collaborative work with Arm’s AI Platforms team and world-class technical teams.
Arm

About Arm

5,001-10,000 employees

Arm’s foundational technology is defining the future of computing. A future built by the greatest technology ecosystem in the world. A future built on Arm. Arm is everywhere technology matters. Technology matters everywhere. Together, we’ll power every technology revolution moving forward, including cloud computing, automotive and autonomous systems, IoT, the metaverse, and beyond. Changing the world. Again. On Arm.

Contact me