Nvidia

Senior Storage Production Engineer - DGX Cloud

Nvidia
Apply
2 hours ago
Remote, AustraliaSenior

Responsibilities

  • Design, implement, and support scalable, highly available storage clusters while maintaining data integrity.
  • Develop storage monitoring, logging, alerting, fault detection, and remediation systems.
  • Optimize storage architectures for AI/ML workloads through low-latency access, caching, capacity management, and performance tuning.
  • Manage the storage service lifecycle from design and deployment through operation, launch reviews, and continuous optimization.
  • Maintain production storage infrastructure by monitoring availability, latency, capacity, and system health.
  • Improve storage efficiency using compression, deduplication, tiering, replication, erasure coding, and intelligent workload placement.
  • Implement automated storage migration, backup, disaster recovery, encryption, access controls, auditing, and compliance mechanisms.
  • Participate in incident response, blameless root cause analysis, and an on-call rotation.

Requirements

  • Bachelor’s degree or equivalent experience in Computer Science, Storage Systems, or a related technical field, with 8+ years of practical experience.
  • Experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise storage systems.
  • Strong understanding of block, file, and object storage scalability, reliability, performance, and operational processes.
  • Experience with NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Expertise in algorithms, data structures, complexity analysis, software design, and automation of large-scale Linux-based storage systems.
  • Experience with one or more of C/C++, Java, Python, Go, NodeJS, and Bash for storage automation, monitoring, and performance tuning.
  • Hands-on experience with Ansible, Chef, Puppet, and Terraform for infrastructure configuration and storage deployment automation.
  • Experience with InfluxDB, Prometheus, Grafana, and the Elastic stack for observability, monitoring, and tracing.
  • Preferred experience includes storage capacity planning, performance tuning, troubleshooting, replication, erasure coding, Git, code review, pipelines, CI/CD, infrastructure as code, and distributed-system performance analysis.
  • Preferred experience includes Kubernetes, OpenStack, private and public cloud storage, hybrid cloud architectures, automated migration, backup, and disaster recovery.

Tech Stack

AnsibleBashCC++ChefGitGoGrafanaInfluxDBJavaKubernetesLinuxOpenStackPrometheusPuppetPythonTerraform

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me