2 months ago
Prague, CzechiaSenior
Responsibilities
- Build a distributed system capable of running millions and eventually billions of AI agents on the platform.
- Build an orchestrator that places sandboxes on appropriate nodes and supports sandbox live migrations.
- Improve the self-hosting developer experience for the open-source platform.
- Keep sandbox startup time below 200 milliseconds from user action.
- Scale to millions and later billions of concurrently running sandboxes.
- Build an observability stack starting at the virtual-machine kernel level.
- Optimize infrastructure performance and efficiency across systems, virtualization, scheduling, networking, and multi-tenant isolation.
Requirements
- At least 5 years of experience building distributed systems and operating infrastructure at serious scale, including 100K+ RPS, multi-region systems, or PB-scale data.
- Deep Linux internals expertise, including kernel-level debugging with eBPF, CPU scheduling, memory management, and cgroups v1 and v2.
- Experience with VM hypervisors such as Firecracker, QEMU, or KVM, including virtio, hypercalls, and nested virtualization trade-offs.
- Strong systems programming skills in at least one of Go, Rust, C, or C++, including performance-critical code and technologies such as lock-free data structures, memory-mapped files, or io_uring.
- Production orchestration experience with Kubernetes, Nomad, or custom orchestration systems, including bin-packing, resource scheduling, and noisy-neighbor problems.
- Experience profiling production systems under load, optimizing hot paths, and working with CPU caches, memory locality, and p99 latency.
- Strong networking knowledge covering L4/L7 load balancing, network namespaces, iptables/nftables, and secure isolated multi-tenant network topologies.
- Comfort working on open-source code and infrastructure and contributing documentation and community discussions.
- Preferred experience with userfaultfd, copy-on-write, lazy loading, GPU passthrough, PCIe device virtualization, AI/ML infrastructure, Firecracker or Cloud Hypervisor open-source contributions, or observability at scale.
Benefits
- In-person work with 4 days on-site and 1 day working from home in San Francisco or Prague, Czech Republic
- Full healthcare, vision, and dental insurance
- Unlimited paid time off
