
Software Engineer, ML Infra
Thinking Machines Lab4 hours ago
Base Salary
$350k - $475k/yr
Responsibilities
- Debug root causes across the kernel, NCCL, scheduler, application, and telemetry layers.
- Provide embedded, hands-on support during critical training runs and major incidents through resolution.
- Act as the primary technical contact for researchers when infrastructure problems have unclear ownership.
- Lead postmortems and build tooling to prevent recurring incidents.
- Mentor engineers in cross-stack infrastructure debugging and operations.
Requirements
- Demonstrated hands-on competence in at least four of Linux kernel, networking, GPUs/CUDA, distributed systems runtimes, storage, compilers/language runtimes, and observability internals.
- Ability to operate without clearly defined ownership and exercise sound escalation judgment.
- Preferred: experience as an escalation point across two or more companies.
- Preferred: meaningful contributions across three or more distinct technical stacks.
- Preferred: experience leading major incidents with non-obvious root causes.
- Preferred: comfort working directly with researchers and communicating workload constraints clearly.
- Preferred: experience operating frontier training or inference clusters and owning difficult, undefined infrastructure problems.
Benefits
- Annual base salary of $350,000–$475,000 USD plus equity.
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- Role is based in San Francisco, California.