
Thinking Machines Lab
Open Positions at Thinking Machines Lab
17 open positions
Build the distributed data infrastructure powering large-scale multimodal datasets and LLM research. You’ll work with researchers to create reliable ingestion, processing, storage, quality, and search systems across petabytes of data.
Own and evolve the security infrastructure behind foundation models, spanning cloud platforms, Kubernetes, identity, data systems, and secure automation. This role combines infrastructure engineering with security architecture to help research and product teams move quickly while protecting models, data, and environments.
Own the reliability of Tinker’s AI fine-tuning platform, from observability and incident response to distributed training resilience and multi-tenant isolation. This role combines production infrastructure engineering with deep reliability work for large-scale GPU workloads.
Build security into AI products, development workflows, and production systems at Thinking Machines. You’ll partner with product and research teams to develop security controls, automation, and defenses against emerging AI-specific risks.
Build and optimize the infrastructure that enables fast, reliable, and efficient inference for large AI models. You’ll collaborate with researchers and engineers on distributed serving, GPU utilization, and scalable AI systems.
Build the platforms, AI coding tools, and secure development environments that make software engineering faster and more reliable across the company. You’ll combine developer enablement with hands-on systems, tooling, and AI integration work.
Build the distributed systems and infrastructure that power frontier-model training, research, and AI products. This generalist role spans core infrastructure, data platforms, and developer productivity, with opportunities to work directly with researchers and engineering teams.
Build and scale AI-focused products across Python and Rust backends, React and TypeScript frontends, and reliable production systems. This evergreen full-stack role offers broad ownership from prototype to launch and ongoing product improvement.
Build the numerical foundations and distributed infrastructure that make trillion-parameter model training stable, scalable, and fast. This role combines research in low-precision computation with hands-on systems engineering across kernels, communication frameworks, and training orchestration.
Build the infrastructure that makes large-scale reinforcement learning and model post-training reliable, efficient, and production-ready. This research engineering role combines deep learning systems, distributed infrastructure, observability, and close collaboration with AI researchers.
Unlock 7 More Matching Jobs
Choose a plan to access all jobs matching your criteria.
First to know
Discover the latest jobs before everyone else
Zero Spam
No ghost jobs, reposts, or sponsored listings
AI-powered filters
Find the most relevant jobs for you