4 hours ago
Base Salary
$231k - $342k/yr
Responsibilities
- Own the technical health and outcomes of assigned AI cloud accounts through handoff, steady state, expansion, and renewal.
- Map customer training, fine-tuning, and inference workloads, including frameworks, schedulers, parallelism, data paths, and performance baselines.
- Lead customer POCs by defining success criteria, coordinating provisioning and capacity, overseeing tests, and driving clear workload decisions.
- Design and defend reference architectures across compute, networking, storage, connectivity, and scheduler integration.
- Develop SLA uptime, downtime, and service-credit methodologies and validate breach events to the root-cause level.
- Build dashboards, telemetry, uptime histories, health scores, and churn early-warning tooling for customer and account health.
- Lead technical escalations and high-severity incidents, including root-cause analysis, RCAs, engineering coordination, and proactive communications.
- Represent the technical voice of the customer through feature-request pipelines, hands-on platform audits, and product-led POCs.
- Track GPU roadmaps, competing clouds, neoclouds, and evolving training and inference technologies to inform customers and internal teams.
Requirements
- 5+ years of experience in technical account management, solutions engineering or architecture, ML engineering, technical program or product management, or infrastructure engineering with significant customer-facing scope in cloud, HPC, or AI infrastructure.
- Hands-on fluency provisioning, benchmarking, and debugging GPU infrastructure across compute, InfiniBand and Ethernet networking, storage, and Slurm or Kubernetes schedulers.
- Working command of AI/ML training, fine-tuning, and inference workloads sufficient to lead technical discussions with ML and infrastructure engineers.
- Experience leading structured technical engagements such as POCs, architecture designs, benchmark programs, or high-severity escalations.
- Experience with scripting, SQL, dashboarding, and turning operational data into dependable tools.
- Ability to communicate deeply technical content effectively in writing and in executive-facing settings.
- Ability to operate amid ambiguity and create methodology where none exists.
- Preferred experience with AI clouds, neoclouds, hyperscalers, large-scale GPU or HPC customers, applied LLM workloads, or the NVIDIA ecosystem.
- Preferred familiarity with NVIDIA NIM, NeMo, Base Command/BCM, CUDA, DGX-class systems, SLA structures, service credits, enterprise contracts, NRR/GRR, product management, or technical program management.
Benefits
- The role requires working from the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday currently designated as the work-from-home day.
- Generous cash and equity compensation are offered, with no specific salary amount stated.
- Health, dental, and vision coverage are provided for employees and dependents.
- Wellness and commuter stipends are available for select roles.
- 401(k) plan with a 2% company match is available to U.S. employees.
- Flexible paid time off plan.
Tech Stack
Categories
Solutions Engineering
