DDN

Platform Support Architect

DDN
Apply
2 months ago
Remote, United KingdomSenior

Responsibilities

  • Serve as the primary NVIDIA AI Enterprise and vector database expert for HyperPOD customer environments.
  • Own end-to-end troubleshooting across GPUs, NVIDIA AI Enterprise services, vector databases, Kubernetes, containers, networking, and Infinia storage.
  • Diagnose performance issues in RAG and agentic AI workflows, including retrieval, GPU utilization, prompt configuration, and data access.
  • Collect and interpret logs and telemetry, build minimal reproductions, and create high-quality vendor and engineering defect reports.
  • Author support runbooks, triage checklists, diagnostic bundles, implementation guides, tuning checklists, and best-practice playbooks.
  • Collaborate on Prometheus, Grafana, ELK, NetQ, and equivalent dashboards for platform and AI service metrics.
  • Build hands-on labs and proof-of-concepts that validate customer RAG and agentic AI use cases on HyperPOD.
  • Provide field feedback to Product Management and Engineering on compatibility, upgrades, rollback, and observability.
  • Collaborate with NVIDIA solutions architects, OEM architects, Professional Services, and Support teams on reference architectures and support practices.

Requirements

  • At least 5 years of experience in Linux-based infrastructure, SRE, MLOps, platform engineering, or L2/L3 production support; 8+ years of total technical experience is preferred.
  • Strong hands-on experience with containers and Kubernetes, including Docker or containerd, Helm, Operators, pods, DaemonSets, CSI, CNI, and ingress or load balancers.
  • Production experience operating NVIDIA GPU-accelerated workloads, including NVIDIA drivers, CUDA concepts, GPU performance, and NVIDIA GPU Operator.
  • Familiarity with DGX, HGX, or similar GPU cluster platforms.
  • Experience with high-performance storage such as EXAScaler/Lustre, GPFS, Ceph, distributed object storage, enterprise NAS, or SAN.
  • Experience with RDMA-accelerated or high-speed Ethernet/InfiniBand networking, fabrics, switch topologies, and large-scale deployments.
  • Experience with hybrid-cloud or cloud-adjacent Kubernetes and data-locality patterns.
  • Experience operating vector databases such as Milvus, Qdrant, Pinecone, pgVector, OpenSearch, or Elasticsearch vectors.
  • Understanding of RAG and generative AI workflows, including embeddings, retrieval, reranking, prompt design, context management, vector search, and GPU inference.
  • Familiarity with NVIDIA AI Enterprise, NIM, NeMo, NeMo Retriever, NeMo Curator, Triton Inference Server, TensorRT, TensorRT-LLM, CUDA libraries, and NVIDIA enterprise AI blueprints.
  • Experience designing, operating, or supporting MLOps or GenAI pipelines, including model CI/CD, deployment, canarying, rollback, GPU resource management, monitoring, and alerting.
  • Strong diagnostic, communication, stakeholder management, and technical-asset development skills.
  • Preferred qualifications include scale-out GPU/AI storage, RDMA-enabled HPC/AI clusters, NVIDIA reference blueprints, AI observability, responsible AI, regulatory awareness, and observability stacks tuned for AI workloads.

Tech Stack

DockerElasticsearchGrafanaHelmKubernetesLinuxNimPrometheus

Categories

Solutions Engineering
DDN

About DDN

1,001-5,000 employees
Contact me