4 months ago
Remote, United States or Remote, EMEASenior
Responsibilities
- Profile and analyze GPU performance at the system and kernel levels with hardware and development teams.
- Evaluate and compare GPU performance across platforms, architectures, and software stacks including CUDA and ROCm.
- Debug and optimize machine-learning workloads on GPU hardware and resolve performance bottlenecks.
- Conduct acceptance testing for new GPU clusters to verify performance, stability, and compatibility for AI workloads.
- Experiment with GPU system configurations, interconnect strategies, and system-level optimizations to assess performance and scalability.
- Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends.
- Contribute to internal tooling, frameworks, and engineering best practices.
Requirements
- Profound understanding of machine-learning theoretical foundations.
- Deep understanding of performance considerations for large neural-network training and inference, including parallelism, offloading, custom kernels, hardware features, attention optimizations, and dynamic batching.
- Deep experience with PyTorch, JAX, Megatron-LM, and TensorRT-LLM.
- Good understanding of CUDA, NCCL, drivers, and relevant GPU libraries.
- Familiarity with Docker and Kubernetes.
- Strong communication skills and ability to work independently.
- Familiarity with vLLM, SGLang, and TensorRT is preferred.
- Experience with Python and performance-profiling tools such as Nsight, nvprof, and perf is preferred.
- Familiarity with AWS, GCP, and Azure ML is preferred.
- Contributions to open-source machine-learning benchmarking tools are preferred.
- Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.
Benefits
- Competitive compensation (amount not stated)
- Career growth and learning opportunities
- Flexible work-life balance
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Categories
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
