
Senior Systems Engineer - AI Infrastructure
Clockwork.io4 months ago
Base Salary
$150k - $230k/yr
Responsibilities
- Design and implement low-level systems software for GPU clusters.
- Modify and extend PyTorch, NCCL, and the CUDA runtime.
- Build components that improve the reliability and efficiency of large-scale GPU training.
- Debug complex distributed and concurrent systems with subtle, nondeterministic failures.
- Own systems end-to-end from design through production.
- Lead the design of significant system components, define technical direction, and mentor engineers.
Requirements
- Experience designing and building complex systems rather than only deploying or operating them.
- Strong C/C++ skills in systems contexts.
- Deep understanding of concurrency, memory models, and failure modes.
- Ability to reason about consistency, ordering, and partial failures in distributed systems.
- Comfort reading and modifying large, unfamiliar codebases.
- Preferred experience with GPU programming or CUDA, GPU systems, RDMA, InfiniBand, ML framework or runtime internals, and cluster scheduling or orchestration systems.
Benefits
- Competitive compensation and a great benefits package.
- Catered lunch.
- Eligibility to participate in the company equity program, including potential stock options.
- Equal opportunity workplace committed to inclusion.
About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io