
AI Training Infrastructure Engineer
Designworks Talent2 months ago
Bellevue, WA, USASenior / Staff+
Responsibilities
- Build and scale distributed training infrastructure for large AI models across large GPU clusters.
- Design systems that improve training reliability, efficiency, and resource utilization.
- Develop fault-tolerance, checkpointing, recovery, and large-scale training operations solutions.
- Integrate AI models into production training pipelines with platform, orchestration, and performance engineering teams.
- Diagnose and resolve issues affecting training throughput, stability, reliability, and cost efficiency.
- Build tools and automation that improve the developer experience for AI researchers and engineers.
- Establish best practices for training infrastructure, operational processes, and platform reliability.
- Help evolve the AI infrastructure platform as an early member of the engineering team.
Requirements
- Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
- Experience supporting large AI models, foundation models, post-training workflows, or similar machine learning systems.
- Strong understanding of reliability, scalability, and efficiency challenges in multi-node GPU training.
- Experience integrating training systems with production machine learning pipelines.
- Strong programming skills and experience with complex distributed systems.
- Ability to independently own technically challenging projects in a fast-moving engineering environment.
- Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies is preferred.
- Experience with supervised fine-tuning, reinforcement learning from human feedback, or other post-training workflows is preferred.
- Experience operating AI training infrastructure at scale in a hyperscaler, AI research organization, cloud provider, or GPU cloud environment is preferred.
- Experience optimizing GPU utilization, training performance, or distributed system reliability is preferred.
- Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms is preferred.
- U.S. work authorization is required; visa sponsorship is not currently available.
Benefits
- Hybrid role in the Bellevue, WA area with approximately three days per week in the office.
- Relocation encouraged for candidates elsewhere in the U.S.
- Medical, dental, and vision insurance for U.S.-based employees.
- 401(k) plan with company match.
- Paid holidays.
- Some roles may be eligible for merit increases, annual bonuses, and stock.
Tech Stack
Categories
About Designworks Talent
Designworks Talent specializes in the design and delivery of enterprise workforce solutions including strategy, talent acquisition, engagement, and succession planning. We work collaboratively with business leaders and enterprise Talent Acquisition and Human Resources teams to deliver custom workforce solutions to meet the dynamic talent demands of your business.