24 hours ago
Base Salary
$209k - $283k/yr
Responsibilities
- Define and build architecture for cloud-based AI inference services.
- Develop Kubernetes controllers and platform capabilities for workload deployment, scheduling, recovery, scaling, upgrades, and lifecycle management.
- Establish production practices for health validation, progressive rollout, rollback, observability, and service objectives.
- Improve platform reliability, scalability, performance, and resource efficiency.
- Lead production readiness reviews and resolve complex issues across services, Kubernetes, networking, and compute infrastructure.
- Turn incidents and operational bottlenecks into durable platform improvements.
- Lead design and build reviews, mentor engineers, and drive technical alignment across teams.
Requirements
- At least 5 years of experience, or equivalent proven impact, building distributed systems, cloud platforms, or production infrastructure.
- Deep production experience with Kubernetes, including controllers, operators, scheduling, resource management, networking, and workload lifecycle management.
- Strong software and production engineering skills, including hands-on programming in Go, C++, Rust, Python, or a similar language.
- Experience with reliable services, APIs, concurrency, observability, deployment safety, capacity planning, and incident response.
- A track record of leading complex technical initiatives while remaining hands-on across architecture, implementation, debugging, mentoring, and cross-team technical direction.
- Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.
- Preferred: experience with AI infrastructure, model serving, or accelerator-backed workloads.
- Preferred: familiarity with PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.
- Preferred: knowledge of inference performance and resource-efficiency considerations.
Benefits
- Relocation package, including visa sponsorship support, is available for candidates who require it.
- Hybrid working arrangements are determined by the team, with role-specific details provided during application.
- Competitive and equitable total reward package, with details shared during recruitment.
- Accommodation and adjustment support is available during the recruitment process.
About Arm
Arm’s foundational technology is defining the future of computing. A future built by the greatest technology ecosystem in the world. A future built on Arm. Arm is everywhere technology matters. Technology matters everywhere. Together, we’ll power every technology revolution moving forward, including cloud computing, automotive and autonomous systems, IoT, the metaverse, and beyond. Changing the world. Again. On Arm.
