Nvidia

Senior Solutions Architect, DevOps

Nvidia
Apply
20 hours ago
Tel Aviv-Yafo, IsraelSenior

Responsibilities

  • Advise on and help maintain large-scale computational and AI infrastructure, including monitoring, logging, workload orchestration, Kubernetes, and Linux job schedulers.
  • Provide consultative guidance and hands-on troubleshooting across bare metal, operating systems, software stacks, container platforms, networking, and storage.
  • Assess customer environments and recommend production-ready Kubernetes container platforms integrated with enterprise networking and storage.
  • Develop and document standard methodologies, operational guidelines, runbooks, onboarding materials, and best-practice guides.
  • Support development activities and conduct proofs of concept and proofs of value for new features, architectures, and upgrade approaches.
  • Lead technical engagements for assigned customer accounts and influence long-term DevOps, platform architecture, infrastructure, and operations decisions.
  • Lead architectural reviews and present scalable infrastructure solutions to customer and executive stakeholders.

Requirements

  • Bachelor’s, master’s, or PhD in computer science, electrical or computer engineering, physics, mathematics, or a related field.
  • 5+ years of professional experience managing scalable cloud environments and working in automation engineering roles.
  • Proven networking, data center architecture, and HPC/AI cluster deployment, optimization, and troubleshooting experience.
  • Hands-on experience deploying, configuring, and optimizing NVIDIA GPU-accelerated infrastructure, including driver management, CUDA integration, and GPU workload profiling.
  • Extensive Kubernetes experience covering container orchestration, resource scheduling, scaling, and integration with GPU-accelerated and HPC environments.
  • Strong familiarity with HPC and AI technologies, including CPUs, GPUs, high-speed interconnects, and supporting software stacks.
  • Deep knowledge of Linux, including Red Hat and Ubuntu, OS-level security, and relevant protocols.
  • Proficiency with Python, Bash, configuration management, and infrastructure-as-code tools such as Ansible and Terraform.
  • Experience with Grafana, Loki, and Prometheus for monitoring, logging, observability, and fault-tolerant systems.
  • Strong solution architecture, customer consulting, architectural review, and executive presentation skills.
  • Preferred experience with CI/CD pipelines, Kubernetes operators for GPU and network management, SLURM, MPI, enroot, job provisioning, cluster change management, NVIDIA Base Command Manager, and RDMA fabrics such as InfiniBand or RoCE.

Tech Stack

AnsibleBashGrafanaKubernetesLinuxPrometheusPythonTerraform

Categories

Solutions Engineering
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me