Member of Technical Staff, Training Infra
The Inception Company6 months ago
San Mateo, CA, USASenior
Responsibilities
- Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
- Develop high-performance optimizations to maximize training throughput and efficiency.
- Build reusable frameworks and libraries that improve training reproducibility, reliability, and scalability for new model architectures.
- Debug and maintain performant, maintainable code in complex codebases.
Requirements
- BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
- Understanding of PyTorch and TensorFlow from a systems perspective.
- Strong engineering skills with the ability to contribute performant, maintainable code and debug complex codebases.
- Proficiency in Python and at least one systems programming language: C++, Rust, or Go.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
- Preferred: experience building and maintaining language models with tens of billions of parameters or more.
- Preferred: experience with Kubeflow or Airflow.
- Preferred: background in performance optimization and profiling of ML systems using Prometheus, Grafana, or OpenTelemetry.
- Preferred: familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM.