
Software Engineer, Systems Generalist
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Architect and scale the core infrastructure supporting foundation-model training, research, and product development.
- Build and operate reliable Kubernetes clusters with GPU workloads and infrastructure supporting Tinker.
- Design, optimize, and maintain scalable data pipelines and data infrastructure using Spark and other modern data infrastructure technologies.
- Build tooling, systems, and frameworks that provide well-configured and optimized developer environments.
- Work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable insights across models, products, and data assets.
- Contribute across infrastructure teams according to organizational needs, interests, and experience.
Requirements
- Bachelor’s degree in computer science, engineering, or a similar field, or equivalent experience.
- Proficiency in at least one backend language, particularly Python or Rust.
- Experience operating large-scale clusters and container orchestration systems such as Kubernetes or Slurm.
- Ability to work across the stack and own projects end to end.
- Preferred: strong debugging across application, operating-system, and network layers.
- Preferred: proficiency in Python or Rust, containers, and modern CI.
- Preferred: experience with Kubernetes, controllers or operators, or performance profiling.
- Preferred: familiarity with GPU/ML workflows or large-scale data/evaluation pipelines.
- Ability to collaborate with cross-functional partners and subject-matter experts and take initiative across teams and technology stacks.
Benefits
- Generous health, dental, and vision benefits; unlimited PTO; paid parental leave; and relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.