GrepJob
Designworks Talent

AI Training Infrastructure Engineer

Designworks Talent
Apply
about 5 hours ago
Bellevue, WA, USASenior / Staff+

Responsibilities

  • Build and scale distributed training infrastructure for large AI models across GPU clusters.
  • Design and improve systems for training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, and large-scale training operations.
  • Integrate AI models into production training pipelines with engineering teams.
  • Diagnose and resolve issues affecting training throughput and reliability.
  • Build tools and automation to enhance the developer experience for AI researchers.
  • Establish best practices for training infrastructure and operational processes.
  • Contribute to the evolution of the AI infrastructure platform.

Requirements

  • Hands-on experience with distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models and post-training workflows.
  • Strong understanding of reliability and efficiency challenges in multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience with complex distributed systems.
  • Ability to independently manage technically challenging projects in a fast-paced environment.

Benefits

  • Competitive base pay for the Bellevue market.
  • Eligibility for merit increases, annual bonuses, and stock based on performance.
  • Access to medical, dental, and vision insurance.
  • 401(k) plan with company match.
  • Paid holidays each calendar year.

Categories

AI & MLData Engineering