
Member of Technical Staff, AI Training Infrastructure
Fireworks AI9 days ago
Base Salary
$200k - $350k/yr
Responsibilities
- Design and implement scalable infrastructure for large-scale model training workloads.
- Develop and maintain distributed training pipelines for LLMs and multimodal models.
- Optimize training performance across multiple GPUs, nodes, and data centers.
- Implement monitoring, logging, and debugging tools for training operations.
- Architect and maintain data storage solutions for large-scale training datasets.
- Automate infrastructure provisioning, scaling, and orchestration for model training.
- Collaborate with researchers to implement and optimize training methodologies.
- Analyze and improve the efficiency, scalability, and cost-effectiveness of training systems.
- Troubleshoot complex performance issues in distributed training environments.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- At least 3 years of experience with distributed systems and ML infrastructure.
- Experience with PyTorch.
- Proficiency with AWS, GCP, or Azure.
- Experience with Kubernetes and Docker.
- Knowledge of distributed training techniques including data parallelism, model parallelism, and FSDP.
- Preferred: master’s or PhD in Computer Science or a related field.
- Preferred: experience training large language models or multimodal AI systems.
- Preferred: experience with ML workflow orchestration tools and high-performance distributed computing optimization.
- Preferred: familiarity with ML DevOps practices and contributions to open-source ML infrastructure or related projects.
Benefits
- Opportunity to work on large-scale AI infrastructure and production model training.
- Collaboration with world-class engineers, AI researchers, and an ambitious team.
- Work on bleeding-edge technology with significant ownership and impact.
- Inclusive, equal-opportunity workplace.