AMD

Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD
Apply
23 hours ago
San Jose, CA, USASenior

Responsibilities

  • Design and develop Fleet Manager distributed control-plane services, APIs, schedulers, inference gateways, and command-line tools.
  • Build orchestration and scheduling capabilities for GPU training, inference, custom jobs, and interactive development workloads.
  • Implement reconciliation, lifecycle management, retries, idempotency, recovery, and PostgreSQL-backed data management.
  • Integrate with Kubernetes, Kueue, JobSet, container runtimes, storage systems, observability platforms, Slurm, and Spur.
  • Develop GPU health, diagnostics, quarantine, remediation, secure multi-tenant infrastructure, and inference reliability capabilities.
  • Maintain stable APIs, data models, compatibility contracts, operational procedures, and automated unit, integration, failure-injection, and production-readiness tests.
  • Diagnose complex failures across distributed services, networking, storage, GPU runtimes, drivers, and hardware.
  • Provide technical leadership through architecture and design reviews, code reviews, mentoring, and cross-functional problem solving.

Requirements

  • Strong systems-software development experience in Rust, C++, Go, or a comparable language; production Rust experience is highly desirable.
  • Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.
  • Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • Experience with reliable services using streaming, WebSocket, or gRPC APIs.
  • Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.
  • Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, optimistic concurrency, command-line tools, and stable APIs.
  • Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.
  • Experience with production observability, source control, continuous integration, automated testing, profiling, and debugging tools.
  • Ability to lead technically challenging projects, collaborate across organizational boundaries, and communicate effectively in writing and verbally.
  • Beneficial experience includes AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, GPU device plugins, distributed AI training, multi-node communication, vLLM, SGLang, PyTorch, GPU topology, Kubernetes networking and storage, inference routing, shared storage, GPU infrastructure, and production security.
  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.

Benefits

  • AMD benefits are referenced through the company’s benefits-at-a-glance information.
  • The role is hybrid, as indicated by the posting’s #LI-HYBRID designation.
AMD

About AMD

10,000+ employees

AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.

Contact me