
AI Training Infrastructure Engineer
Designworks Talentabout 5 hours ago
Bellevue, WA, USASenior / Staff+
Responsibilities
- Build and scale distributed training infrastructure for large AI models across GPU clusters.
- Design and improve systems for training reliability, efficiency, and resource utilization.
- Develop solutions for fault tolerance, checkpointing, and large-scale training operations.
- Integrate AI models into production training pipelines with engineering teams.
- Diagnose and resolve issues affecting training throughput and reliability.
- Build tools and automation to enhance the developer experience for AI researchers.
- Establish best practices for training infrastructure and operational processes.
- Contribute to the evolution of the AI infrastructure platform.
Requirements
- Hands-on experience with distributed training systems or large-scale machine learning infrastructure.
- Experience supporting large AI models and post-training workflows.
- Strong understanding of reliability and efficiency challenges in multi-node GPU training.
- Experience integrating training systems with production machine learning pipelines.
- Strong programming skills and experience with complex distributed systems.
- Ability to independently manage technically challenging projects in a fast-paced environment.
Benefits
- Competitive base pay for the Bellevue market.
- Eligibility for merit increases, annual bonuses, and stock based on performance.
- Access to medical, dental, and vision insurance.
- 401(k) plan with company match.
- Paid holidays each calendar year.
Tech Stack
Categories
AI & MLData Engineering