AMD

Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD
Apply
23 hours ago
San Jose, CA, USASenior

Responsibilities

  • Design and develop Fleet Manager distributed control-plane services, APIs, schedulers, inference gateways, and command-line tools.
  • Build orchestration for GPU training, inference, custom jobs, and interactive development workloads.
  • Develop scheduling and admission-control capabilities including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node coordination.
  • Implement reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.
  • Integrate with Kubernetes, Kueue, JobSet, container runtimes, storage systems, and observability platforms.
  • Support portable orchestration across Kubernetes, Slurm, Spur, and future execution environments.
  • Develop GPU health, diagnostics, quarantine, and remediation capabilities using ROCm and AMD hardware telemetry.
  • Build secure multi-tenant infrastructure with authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults.
  • Improve AI inference reliability and performance through routing, streaming, load shedding, health detection, and usage metering.
  • Define APIs, data models, compatibility contracts, operational procedures, and automated unit, integration, failure-injection, and production-readiness tests.
  • Diagnose distributed failures across services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware.
  • Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional collaboration.

Requirements

  • Strong systems-software development experience in Rust, C++, Go, or a comparable language; production Rust is highly desirable.
  • Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.
  • Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • Experience building services using REST, streaming, WebSocket, or gRPC APIs.
  • Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.
  • Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency.
  • Experience developing command-line tools and stable, user-focused APIs.
  • Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.
  • Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis.
  • Experience with source control, continuous integration, automated testing, profiling, and debugging tools.
  • Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries.
  • A bachelor's or master's degree in computer science, computer engineering, electrical engineering, or a related field, or equivalent practical experience.
  • Beneficial experience includes AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, GPU device plugins, distributed AI training, multi-node collective communication, vLLM, SGLang, PyTorch, GPU topology, capacity management, hardware diagnostics, Kubernetes networking and storage, OpenAI-compatible inference APIs, shared storage, GPU infrastructure, and production security or threat modeling.

Benefits

  • Hybrid work arrangement, as indicated by the #LI-HYBRID designation.
  • AMD benefits are offered, with details provided through AMD's benefits-at-a-glance materials.
AMD

About AMD

10,000+ employees

AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.

Contact me