about 13 hours ago
Los Angeles, CA, USA or Seattle, WA, USASenior / Staff+
H1B Sponsor
Responsibilities
- Apply expertise to large-scale compute system development, including design validation and optimization.
- Act as the primary technical point of contact for compute infrastructure across all customer accounts.
- Lead technical implementation, integration testing, and provisioning for GPU/CPU clusters.
- Provide expertise on compute-specific elements including GPUs/CPUs and high-performance networking.
- Handle technical escalations and drive resolution to ensure compliance with SOW and SLA.
- Represent compute systems in cross-functional discussions and issue resolution.
- Interface with various internal teams to ensure successful end-to-end delivery.
- Participate in verification testing and continuous optimization.
- Support expansion planning and capacity scaling across customer accounts.
- Ensure on-time deliverables and proactive risk mitigation for all customers.
Requirements
- Bachelor’s degree in computer science, Electrical Engineering, or a related technical discipline.
- 7+ years of hands-on experience in large-scale compute or data center environments.
- Experience leading technical implementations involving GPUs/CPUs and high-performance systems.
- Master’s degree or higher in a STEM discipline is preferred.
- Prior experience as a Technical Account Manager or Solutions Architect is preferred.
- Deep knowledge of InfiniBand, RoCE, and NVIDIA networking technologies is preferred.
- Experience with AI/ML training and inference infrastructure at scale is preferred.
- Strong background in performance tuning and hybrid/cloud integration is preferred.
- Excellent written and verbal communication skills are required.
- Hands-on experience with Linux environments and infrastructure monitoring is preferred.