20 hours ago
Zürich, SwitzerlandSenior
Responsibilities
- Design, implement, and maintain large-scale HPC and AI clusters with monitoring, logging, and alerting.
- Manage Linux workloads, job scheduling, and orchestration tools.
- Develop and maintain continuous integration and delivery pipelines.
- Automate deployment, management, monitoring, alerting, and self-service resource consumption for large-scale infrastructure.
- Deploy monitoring solutions across servers, networks, and storage.
- Troubleshoot systems from bare metal and operating systems through software stacks and applications.
- Develop and document standard methodologies for internal teams.
- Support research and development activities and participate in proofs of concept and proofs of value.
Requirements
- Degree in Computer Science, Engineering, or a related field and 8+ years of experience.
- Knowledge of HPC and AI solution technologies spanning CPUs, GPUs, high-speed interconnects, and supporting software.
- Experience with workload scheduling and orchestration tools such as Slurm and Kubernetes.
- Strong knowledge of Windows, Linux distributions, networking, operating-system internals, ACLs, OS security protections, and common network protocols.
- Experience with storage solutions such as Lustre, GPFS, and Weka.io, plus familiarity with emerging storage technologies.
- Experience with Python programming and Bash scripting.
- Experience with automation and configuration-management tools such as Jenkins, Ansible, Puppet, and Chef.
- Deep knowledge of networking protocols including InfiniBand and Ethernet.
- Strong experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix.
- Familiarity with cloud platforms including AWS, Azure, and Google Cloud.
- Preferred experience includes CPU or GPU architecture, Kubernetes and container-related technologies, GPU-focused hardware and software such as DGX and CUDA, and RDMA fabrics including InfiniBand or RoCE.
Tech Stack
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
