1 day ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Design and evolve scalable architectures for multi-node LLM inference across GPU clusters.
- Develop infrastructure to optimize production model-serving latency, throughput, and cost efficiency.
- Collaborate with model, systems, compiler, networking, and partner teams on complete high-performance solutions.
- Prototype KV-cache handling, tensor and pipeline parallel execution, and dynamic batching approaches.
- Evaluate and integrate software and hardware technologies related to Spectrum-X, including load balancing, telemetry, congestion control, and vertical application integration.
- Translate high-level architecture into reliable, high-performance systems with internal teams and external partners.
- Author design documents, internal specifications, and technical blog posts, and contribute to open-source efforts when appropriate.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science, Electrical Engineering, or equivalent experience.
- 8+ years of experience building large-scale distributed systems or performance-critical software.
- Deep understanding of deep learning systems, GPU acceleration, AI model execution flows, and/or high-performance networking.
- Strong software engineering skills in C++ and/or Python, preferably with familiarity with CUDA or similar platforms.
- Strong system-level understanding across memory, networking, scheduling, and compute orchestration.
- Excellent communication and collaboration skills across diverse technical domains.
- Preferred experience with LLM training or inference pipelines, transformer model optimization, model-parallel deployments, performance profiling, distributed communication patterns, congestion control, load balancing, and complex organizational processes.
Benefits
- NVIDIA offers highly competitive salaries and a comprehensive benefits package for employees and their families.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
