1 day ago
Base Salary
$224k - $431k/yr
Responsibilities
- Build and manage core attestation cloud services, APIs, SDK and CLI integration points.
- Convert hardware trust mechanisms and standards into production-ready solutions for NVIDIA platforms.
- Improve reliability through SLOs/SLIs, alerting, runbooks, incident response, safe rollouts, and operational improvements.
- Design resilient service behavior for dependency failures, caching challenges, regional issues, and customer-side resilience needs.
- Architect distribution of certificate status, revocation updates, signed metadata, RIM artifacts, trust bundles, and offline verification materials.
- Implement appraisal, policy, and verification workflows for attestation evidence, endorsements, reference values, certificate status, and security requirements.
- Provide technical leadership and mentoring in distributed systems, production debugging, reliability, observability, and operational simplicity.
Requirements
- BS or MS in Computer Science, Information Security, or a related field, or equivalent experience.
- 12+ years of experience designing and building large-scale distributed systems or cloud services, including at least 3 years in security, attestation, or trusted computing.
- Experience owning production services, including monitoring, alerting, incident response, root-cause analysis, and long-term operational improvements.
- Experience developing and maintaining REST and/or gRPC APIs, microservices, control planes, background tasks, caches, queues, data stores, and customer-facing integrations.
- Proficiency in Go or Java; C++ or Rust experience is a plus.
- Hands-on experience with Kubernetes, containers, CI/CD, infrastructure as code, observability tools, and at least one major cloud provider such as AWS, GCP, Azure, or OCI.
- Knowledge of distributed-systems practices including retries, timeouts, circuit breakers, backpressure, rate limiting, caching, consistency, idempotency, failover, and dependency isolation.
- Understanding of PKI, TLS/mTLS, certificate lifecycle, signing, secrets management, authentication and authorization, and secure service-to-service communication.
- Preferred experience with security-critical systems, attestation, confidential computing, trusted execution environments, TPMs, DICE, SPDM, IETF RATS, EAT, CoRIM, HSMs, hardware roots of trust, secure boot, GPU/accelerator attestation, or related standards.
- Preferred experience operating customer-facing services with formal SLOs, error budgets, safe deployment systems, production-readiness reviews, resilience testing, or chaos/game-day practices.
- Preferred experience with identity, PKI, certificate status and revocation, signing, secrets management, trust-material distribution, control-plane reliability, or multi-region, multi-cloud, sovereign, GovCloud, edge, disconnected, air-gapped, or customer-hosted environments.
Benefits
- Eligible for equity and benefits.
- Remote position, indicated by the #LI-Remote designation.
- Applications are accepted at least until September 25, 2026.
Tech Stack
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
