7 days ago
Responsibilities
- Own customer GPU cluster deployments through bring-up, burn-in, production-readiness certification, workload onboarding, performance validation, and ongoing operation.
- Deploy and tune large-scale training and inference stacks across cloud, NeoCloud, and bare-metal environments.
- Lead root-cause analysis, Sev-1 response, and permanent resolution of production incidents.
- Deploy agentic AI solutions and own their production behavior within customer engagements.
- Build observability, benchmarking, and validation tooling for cluster certification and ongoing production health.
- Transfer operational knowledge through documentation, runbooks, and hands-on enablement to customer teams.
- Contribute field learnings and upstream fixes to ROCm, serving frameworks, reference architectures, and AMD product teams.
Requirements
- At least 5 years of production software or infrastructure engineering experience, including operating or deploying systems built by others.
- Hands-on GPU-computing experience at scale, including cluster deployment, distributed training or high-throughput inference, performance debugging, and workload optimization.
- Working knowledge of Kubernetes and/or Slurm, containerized GPU workloads, RCCL/NCCL, RoCE or InfiniBand, and observability tools such as Prometheus and Grafana.
- Cloud platform experience with AWS, Azure, GCP, or NeoCloud environments, including hybrid and bare-metal deployment patterns.
- Proficiency in Python and at least one systems language, with the ability to navigate and modify large unfamiliar codebases.
- Familiarity with inference serving, RAG, and agentic workflows sufficient to deploy and troubleshoot them in customer environments.
- Direct customer-facing experience through embedded deployments, technical escalations, on-site engagements, or equivalent work.
- Experience with ROCm and AMD Instinct GPUs is strongly preferred; deep CUDA ecosystem experience is also valued.
- Open-source contributions to AI/ML infrastructure projects are a plus.
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
Benefits
- Hybrid work arrangement is indicated in the posting.
- AMD benefits are referenced, with details available through the company’s benefits overview.
Tech Stack
Categories
Forward Deployed
About AMD
We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences – the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our mission is the AMD culture. We push the limits of innovation to solve the world’s most important challenges. We strive for execution excellence while being direct, humble, collaborative, and inclusive of diverse perspectives. AMD together we advance_