9 hours ago
Markham, CanadaStaff+
Responsibilities
- Own ROCm’s end-to-end validation architecture across functional, workload, performance, stress, stability, scale-out, and system-level test layers.
- Define release-qualification gates, exit criteria, coverage standards, performance baselines, stability targets, scale targets, and RAS criteria.
- Architect distributed test runners, CI fleets, hardware-lab orchestration, result data lakes, flaky-test detection, bisection automation, and developer pre-submit pipelines.
- Establish GitHub-based quality workflows, PR gating policies, required checks, code-coverage standards, bug-bash cadences, and issue-management practices.
- Lead root-cause analysis of complex multi-component, multi-node, hardware, firmware, and customer escalations.
- Drive validation of multi-GPU server nodes, PCIe, Infinity Fabric, xGMI, BMC/IPMI, thermal and power behavior, firmware interactions, and Ethernet/InfiniBand/UALink fabrics.
- Lead AI/ML and HPC workload validation for training, inference, recommender systems, scientific kernels, and benchmark suites.
- Mentor Senior and Staff validation engineers, SDETs, and SQA leads through technical reviews and written guidance.
- Influence validation roadmaps for next-generation Instinct GPUs and represent ROCm validation in customer, OEM, and open-source engagements.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, or a related discipline, or equivalent demonstrated experience.
- Software engineering experience in validation, SDET, or quality engineering, including leadership of complex systems validation.
- Expert Python skills for test automation and infrastructure and strong C++ skills for debugging and production-code extensions.
- Deep expertise in at least two relevant areas, including GPU software stacks, AI/ML frameworks, HPC runtimes and communication libraries, Linux kernel or drivers, accelerator firmware, or distributed systems.
- Experience validating multi-GPU, multi-node server platforms through stress, soak, fault-injection, and RAS testing.
- Experience defining release-qualification programs for hyperscalers, OEMs, or Tier-1 customers.
- Experience contributing to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
- Experience with agentic AI workflows, automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
- Experience validating or operating large-scale GPU clusters of 256 or more GPUs, including fabric bring-up, health monitoring, and diagnostics.
- Familiarity with AI training, inference, HPC benchmark methodologies, performance validation, profiling tools, hardware-lab automation, and pre-silicon or first-silicon accelerator bring-up.
Benefits
- Hybrid role located in San Jose, California.
- AMD benefits are offered; details are provided through AMD’s benefits overview.
About AMD
AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.
