Nvidia

Senior Solutions Architect, HPC and AI

Nvidia
Apply
1 month ago
Berlin, GermanySenior

Responsibilities

  • Collaborate with NVIDIA training framework developers and product teams to help partners adopt the latest features.
  • Deploy, debug, and improve the efficiency of AI workloads on large-scale NVIDIA platforms.
  • Benchmark framework features, analyze performance, and share actionable insights with customers and internal teams.
  • Resolve customer cluster performance and stability issues, identify bottlenecks, and implement solutions.
  • Guide customers in scaling workloads efficiently and reliably on the latest NVIDIA GPUs.
  • Help customers implement advanced resiliency features within AI training pipelines as part of Europe’s Sovereign AI initiative.

Requirements

  • BS, MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering field, or equivalent practical experience.
  • 8+ years of experience with accelerated computing technologies at cluster scale, ideally involving NVIDIA platforms.
  • Strong programming skills in at least one of C, C++, or Python.
  • Practical experience identifying and resolving bottlenecks in large-scale training workloads or parallel applications.
  • Hands-on experience profiling and debugging large parallel applications.
  • Understanding of CPU and GPU architectures, CUDA, parallel filesystems, and high-speed interconnects.
  • Experience with large compute clusters and their scheduling and resource-management mechanisms, such as SLURM or cloud-based clusters.
  • Proficiency with training pipelines and frameworks, including their internal operations and performance characteristics.
  • Preferred experience debugging training pipelines across thousands of GPUs in production.
  • Preferred experience with Nsight Systems, Nsight Compute, NCCL, MPI, and low-level communication libraries.
  • Ability to debug stability issues across parallel applications, training frameworks, runtime libraries, schedulers, and hardware.
  • Understanding of PyTorch, Megatron-LM, NeMo, vLLM, Dynamo, TensorRT-LLM, RedHat Inference Server, or SGLang.

Tech Stack

Categories

Solutions Engineering
Nvidia

About Nvidia

10,000+ employees

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Contact me