GrepJob
Thinking Machines Lab

Software Engineer, Supercomputing

Thinking Machines Lab
Apply
5 days ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Operate and automate large GPU clusters, including provisioning, imaging, and capacity planning.
  • Write software that abstracts cluster management and provides a unified interface for training and inference.
  • Extend Kubernetes, Slurm, or similar orchestration systems for topology-aware placement, preemption, quotas, and fair-share multi-tenancy.
  • Monitor and improve speed, reliability, and error-recovery metrics.
  • Build reliable storage and artifact paths for datasets, checkpoints, and logs with retention and lineage.
  • Partner with researchers to unblock scale runs and advise on parallelism and performance trade-offs.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
  • Proficiency in at least one backend language, particularly Python or Rust.
  • Experience operating large-scale clusters and container orchestration systems such as Kubernetes or Slurm.
  • Ability to work across the stack and own projects end-to-end.
  • Strong collaboration and initiative in cross-functional environments.
  • Strong systems background in Linux, networking, and infrastructure-as-code is preferred.
  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference is preferred.
  • Experience supporting large-scale model training or inference environments is preferred.
  • Understanding of deep learning frameworks such as PyTorch, TensorFlow, or JAX and their underlying system architectures is preferred.

Benefits

  • Health, dental, and vision benefits
  • Unlimited paid time off
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship is available
  • Based in San Francisco, California
  • Evergreen role with ongoing application review

Categories