2 months ago
Remote, United KingdomSenior
Responsibilities
- Serve as the primary NVIDIA AI Enterprise and vector database expert for HyperPOD customer environments.
- Own end-to-end troubleshooting across GPUs, NVIDIA AI Enterprise services, vector databases, Kubernetes, containers, networking, and Infinia storage.
- Diagnose performance issues in RAG and agentic AI workflows, including retrieval, GPU utilization, prompt configuration, and data access.
- Collect and interpret logs and telemetry, build minimal reproductions, and create high-quality vendor and engineering defect reports.
- Author support runbooks, triage checklists, diagnostic bundles, implementation guides, tuning checklists, and best-practice playbooks.
- Collaborate on Prometheus, Grafana, ELK, NetQ, and equivalent dashboards for platform and AI service metrics.
- Build hands-on labs and proof-of-concepts that validate customer RAG and agentic AI use cases on HyperPOD.
- Provide field feedback to Product Management and Engineering on compatibility, upgrades, rollback, and observability.
- Collaborate with NVIDIA solutions architects, OEM architects, Professional Services, and Support teams on reference architectures and support practices.
Requirements
- At least 5 years of experience in Linux-based infrastructure, SRE, MLOps, platform engineering, or L2/L3 production support; 8+ years of total technical experience is preferred.
- Strong hands-on experience with containers and Kubernetes, including Docker or containerd, Helm, Operators, pods, DaemonSets, CSI, CNI, and ingress or load balancers.
- Production experience operating NVIDIA GPU-accelerated workloads, including NVIDIA drivers, CUDA concepts, GPU performance, and NVIDIA GPU Operator.
- Familiarity with DGX, HGX, or similar GPU cluster platforms.
- Experience with high-performance storage such as EXAScaler/Lustre, GPFS, Ceph, distributed object storage, enterprise NAS, or SAN.
- Experience with RDMA-accelerated or high-speed Ethernet/InfiniBand networking, fabrics, switch topologies, and large-scale deployments.
- Experience with hybrid-cloud or cloud-adjacent Kubernetes and data-locality patterns.
- Experience operating vector databases such as Milvus, Qdrant, Pinecone, pgVector, OpenSearch, or Elasticsearch vectors.
- Understanding of RAG and generative AI workflows, including embeddings, retrieval, reranking, prompt design, context management, vector search, and GPU inference.
- Familiarity with NVIDIA AI Enterprise, NIM, NeMo, NeMo Retriever, NeMo Curator, Triton Inference Server, TensorRT, TensorRT-LLM, CUDA libraries, and NVIDIA enterprise AI blueprints.
- Experience designing, operating, or supporting MLOps or GenAI pipelines, including model CI/CD, deployment, canarying, rollback, GPU resource management, monitoring, and alerting.
- Strong diagnostic, communication, stakeholder management, and technical-asset development skills.
- Preferred qualifications include scale-out GPU/AI storage, RDMA-enabled HPC/AI clusters, NVIDIA reference blueprints, AI observability, responsible AI, regulatory awareness, and observability stacks tuned for AI workloads.
Tech Stack
Categories
Solutions Engineering