
Software Engineer, Supercomputing
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Operate and automate large GPU clusters, including provisioning, imaging, and capacity planning.
- Write software that abstracts cluster management and provides a unified interface for training and inference.
- Extend Kubernetes, Slurm, or similar orchestration systems for topology-aware placement, preemption, quotas, and fair-share multi-tenancy.
- Monitor and improve speed, reliability, and error-recovery metrics.
- Build reliable storage and artifact paths for datasets, checkpoints, and logs with retention and lineage.
- Partner with researchers to unblock scale runs and advise on parallelism and performance trade-offs.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
- Proficiency in at least one backend language, particularly Python or Rust.
- Experience operating large-scale clusters and container orchestration systems such as Kubernetes or Slurm.
- Ability to work across the stack and own projects end-to-end.
- Strong collaboration and initiative in cross-functional environments.
- Strong systems background in Linux, networking, and infrastructure-as-code is preferred.
- Familiarity with CUDA/NCCL and performance profiling for distributed training and inference is preferred.
- Experience supporting large-scale model training or inference environments is preferred.
- Understanding of deep learning frameworks such as PyTorch, TensorFlow, or JAX and their underlying system architectures is preferred.
Benefits
- Health, dental, and vision benefits
- Unlimited paid time off
- Paid parental leave
- Relocation support as needed
- Visa sponsorship is available
- Based in San Francisco, California
- Evergreen role with ongoing application review