9 hours ago
Responsibilities
- Develop and maintain product-facing supportability features, scripts, configuration guidance, Kubernetes manifests, Helm charts, and reproducible test cases.
- Build Python-based validators, log collectors, reproduction harnesses, and automation for NVIDIA AI Enterprise deployments across NGC and container orchestrators.
- Contribute code fixes, patches, and pull requests with engineering teams to resolve customer-impacting issues and improve product readiness.
- Support enterprise customers deploying NVIDIA AI Enterprise on Kubernetes-based and containerized production AI platforms across datacenter and cloud environments.
- Own customer issues end-to-end by reproducing them, collecting diagnostics, providing mitigations, and partnering on fixes.
- Create detailed bug reports and RFEs with reproduction steps, environment details, impact analysis, and supporting artifacts.
- Develop customer-facing and internal knowledge bases, runbooks, and deployment guidance.
- Provide engineering assistance during a Sev1 customer outage one weekend per month.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent experience.
- At least 5 years of system software development and troubleshooting experience, ideally including customer-facing work.
- Python programming and scripting skills are required; Bash is expected, while Go or C++ is a plus.
- Strong troubleshooting skills involving networking, concurrency, operating-system concepts, and isolating issues across application, platform, and infrastructure layers.
- Deep understanding of at least two areas including datacenters and servers, distributed systems, virtualization, deep learning frameworks, containers, hybrid cloud, or reliable deployment practices.
- Familiarity with GPU-accelerated AI/ML stacks, NGC containers, CUDA concepts, and inference serving such as Triton.
- Deep Linux knowledge and production Linux troubleshooting experience; Windows knowledge is a plus.
- Professional communication and interpersonal skills with a strong problem-solving orientation.
- Preferred experience deploying NVIDIA AI Enterprise in production across on-premises or cloud environments.
- Preferred production Kubernetes operations experience, including cluster upgrades and control-plane and data-plane failure modes.
- Preferred GPU and cloud performance debugging experience, including profiling and latency or throughput tuning.
- Preferred familiarity with observability and tracing tools and AI coding assistants such as Cursor, Claude Code, or Codex.
Benefits
- The role includes customer-facing technical ownership across cloud and datacenter deployments.
- One weekend per month of on-call engineering assistance is required for Sev1 customer outages.
Tech Stack
Categories
Solutions Engineering
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
