1 day ago
Base Salary
$200k - $322k/yr
Responsibilities
- Define and guide the architecture, interfaces, technical vision, and growth of a key DGX Cloud Data Platform domain.
- Lead complex cross-team technical initiatives by turning ambiguous requirements into architectures, writing essential code, coordinating implementation, and integrating systems into production.
- Architect, implement, and evolve batch and streaming systems for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
- Build shared libraries, workflow and orchestration abstractions, deployment tooling, and engineering standards adopted across teams.
- Lead high-impact production investigations across pipelines, applications, query engines, distributed processing, storage, networks, and cloud services.
- Drive engineering standards for testing, data quality, reconciliation, lineage, observability, service-level objectives, secure identities, least-privilege access, release readiness, and auditable deployments.
- Establish data models, semantics, ownership boundaries, and serving interfaces, including tables, APIs, automation, dashboards, and internal applications.
- Provide technical leadership through architecture reviews, build reviews, mentorship, and resolution of difficult engineering tradeoffs.
Requirements
- 12+ years of relevant industry experience.
- A bachelor’s degree or equivalent experience and a master’s degree or equivalent experience in Computer Science, Engineering, or a related field.
- A sustained record of personally building and operating production software, data platforms, databases, or distributed systems with end-to-end technical ownership.
- Deep hands-on experience with distributed processing, analytical or relational databases, production ETL, change-data capture, streaming or event processing, or backend and cloud systems handling large data volumes.
- Production proficiency in a backend or systems language and deep experience with data-processing and platform libraries or frameworks.
- Strong SQL and data-modeling skills, including query execution, incremental processing, schema evolution, consistency, analytical consumption, idempotency, replay, late-arriving data, partial failure, and cross-system correctness.
- Demonstrated ability to diagnose failures using logs, metrics, traces, query plans, profiles, and controlled experiments and implement durable fixes.
- Strong architectural judgment across reliability, performance, cost, security, compatibility, and maintainability, including experience guiding major migrations or architectural changes.
- Experience establishing automated testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment practices adopted by multiple teams.
- Preferred experience with distributed data processing, lakehouse architectures, distributed streaming or event-driven systems, relational and distributed databases, time-series systems, object storage, cloud infrastructure, container orchestration, workload schedulers, compute or GPU clusters, fleet telemetry, or agentic workflow automation.
Benefits
- Competitive base salary ranging from 200,000 USD to 322,000 USD, plus equity and benefits.
- Comprehensive benefits package for employees and their families.
- Applications will be accepted at least until October 9, 2026.
- This posting is for an existing vacancy.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
