AMD

Principal Software Engineer — AI Performance & Reliability

AMD
Apply
15 hours ago
San Jose, CA, USAStaff+
H1B sponsor

Responsibilities

  • Profile and optimize AI model training and inference workloads.
  • Improve model throughput, latency, memory efficiency, scalability, and reliability.
  • Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • Develop performance tooling, benchmarks, automation, and observability systems.
  • Investigate and resolve complex production issues affecting AI workloads.
  • Collaborate with customers to understand requirements, reproduce issues, and recommend solutions.
  • Translate customer feedback into product and infrastructure improvements.
  • Work with machine learning engineers, systems engineers, hardware teams, and product teams.
  • Document performance findings, technical recommendations, and best practices.

Requirements

  • Strong software engineering skills and experience building production-quality systems.
  • Experience with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Understanding of systems performance concepts including latency, throughput, memory bandwidth, utilization, and distributed communication.
  • Proficiency in Python, C++, or similar systems-oriented programming languages.
  • Experience with PyTorch, TensorFlow, or JAX.
  • Strong debugging and analytical skills across multiple technology-stack layers.
  • Clear written and verbal communication skills.
  • Willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.
  • Preferred experience optimizing large language models, diffusion models, or recommendation models.
  • Preferred experience with GPU, accelerator, or distributed computing environments.
  • Preferred familiarity with ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.
  • Preferred experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.
  • Preferred experience operating AI systems in production, customer-facing technical roles, and designing benchmarks or conducting systematic performance analysis.

Benefits

  • Hybrid work arrangement, with San Jose, California or Bellevue, Washington preferred; other US locations may be considered.
  • AMD benefits are offered, with details provided through AMD’s benefits overview.

Tech Stack

Categories

AMD

About AMD

10,000+ employees

AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.

Contact me