AMD

2027 PhD AI Training Systems and Performance Engineer Intern/Co-Op

AMD
Apply
2 days ago
San Jose, CA, USAIntern

Responsibilities

  • Develop and optimize large-scale AI training and fine-tuning workloads on AMD GPU platforms.
  • Bring up foundation models and training frameworks on AMD hardware and establish reproducible correctness and performance baselines.
  • Profile AI workloads and identify bottlenecks across GPUs, CPUs, memory systems, networking, and communication infrastructure.
  • Develop tooling and workflows for automated training setup, debugging, performance analysis, and optimization using LLM-powered agents and agentic AI techniques.
  • Implement optimization strategies to improve training throughput, GPU utilization, memory efficiency, and scalability across distributed multi-GPU environments.
  • Benchmark and tune AI frameworks, libraries, SDKs, and applications using performance analysis tools.
  • Collaborate with software engineers and architects to evaluate AI models, distributed training techniques, and performance opportunities.
  • Turn successful experiments into reusable workflows, best practices, and software improvements for AI training on AMD hardware.

Requirements

  • Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, Machine Learning, or a related technical discipline.
  • Strong programming experience in Python and/or C++.
  • Hands-on experience implementing, training, and debugging deep learning models with PyTorch, JAX, TensorFlow, vLLM, or SGLang.
  • Experience in distributed training systems, parallelism, GPU performance optimization, AI systems software, high-performance computing, LLM training and fine-tuning, or agentic AI/LLM-powered automation.
  • Understanding of transformer architectures, mixture-of-experts models, and modern LLM training techniques.
  • Experience profiling workloads with tools such as PyTorch Profiler, ROCm Profiler, VTune, Nsight, or similar tools.
  • Familiarity with distributed training technologies and communication libraries such as MPI, NCCL/RCCL, and OpenMP.
  • Understanding of GPU architecture, memory systems, communication bottlenecks, and performance tuning methodologies.
  • Preferred experience identifying and resolving compute, memory, data-loading, or communication bottlenecks in large-scale AI workloads.
  • Experience with ROCm, HIP, Triton, GPU kernel optimization, or AI systems software development is a plus.
  • Publications in AI, machine learning, high-performance computing, computer architecture, or related research areas are a plus.

Benefits

  • Full-time hybrid or onsite work in San Jose or Santa Clara, California, throughout the internship or co-op term.
  • Spring/Summer co-op: January 25, 2027 to August 13, 2027.
  • Summer internship: May 24, 2027 to August 13, 2027 for semester students, or June 21, 2027 to September 10, 2027 for quarter students.
  • Summer/Fall co-op: May 24, 2027 to December 10, 2027 for semester students, or June 21, 2027 to December 10, 2027 for quarter students.
  • Benefits are offered through AMD.

Tech Stack

Categories

AMD

About AMD

10,000+ employees

AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.

Contact me