1 day ago
Base Salary
$168k - $270k/yr
Responsibilities
- Build production observability solutions for agentic AI applications, including metrics, tracing, logging, and alerting.
- Develop deployment pipelines and reliability tools for the Agentic AI Factory model.
- Define SLOs and SLIs with AI application teams and ensure production readiness for new agent deployments.
- Instrument LLM workflows to monitor performance, cost, quality, token usage, multi-step reasoning chains, and tool orchestration.
- Lead incident response, root cause analysis, and reliability improvements for BizApps AI services.
- Address observability gaps such as hallucination detection and orchestration failure tracing.
- Collaborate with platform, data, and application engineering teams to integrate reliability throughout the development lifecycle.
Requirements
- Bachelor’s or master’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
- 8+ years of experience and strong software engineering skills in Python.
- Experience with modern CI/CD practices, Kubernetes, Docker, cloud infrastructure, and observability platforms.
- Experience with Datadog, OpenTelemetry, Grafana, or Prometheus.
- An SRE/DevOps approach involving production ownership, toil automation, and reliability engineering.
- Proven ability to debug complex distributed systems.
- Preferred experience with LangChain, LlamaIndex, Semantic Kernel, agentic AI patterns, and production ML/AI systems.
- Preferred familiarity with AI-specific observability, including inference latency profiling, token economics, and response-quality monitoring.
- Preferred contributions to open-source observability or AI tooling projects.
- Preferred experience with Terraform, Pulumi, and GitOps or equivalent workflows.
Benefits
- Base salary range of 168,000 USD to 270,250 USD, determined by location, experience, and comparable employee pay.
- Eligible for equity and benefits.
- Applications accepted at least until September 12, 2026.
- This posting is for an existing vacancy.
Tech Stack
Categories
DevOpsSite Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
