7 months ago
Remote, Worldwide or Amsterdam, NetherlandsSenior
Responsibilities
- Ensure the reliability, availability, and performance of compute nodes running virtual machines.
- Analyze and debug Linux systems across user space and kernel space.
- Troubleshoot production issues involving CPU, memory, NUMA, cgroups, and scheduling.
- Work hands-on with virtualization and containerization using QEMU/KVM and Linux-native technologies.
- Design and evolve observability capabilities including metrics, logs, traces, alerts, SLIs, and SLOs.
- Lead incident response, root-cause analysis, and postmortems while driving long-term reliability improvements.
- Collaborate with platform, kernel/hypervisor, GPU, and infrastructure teams to improve system design and operability.
Requirements
- Deep expertise in Linux user space, kernel space, kernel subsystems, system boundaries, and system constraints.
- Hands-on experience with QEMU/KVM and understanding of VM lifecycles, performance characteristics, and failure modes.
- Practical experience with containers, namespaces, and cgroups, including resource isolation and control.
- Strong ability to analyze complex system failures using a structured, hypothesis-driven approach.
- Understanding of SRE principles, system design and operations, and experience building and operating observability stacks.
- Experience with Kubernetes internals or node-level components is preferred.
- Experience with perf, eBPF, ftrace, strace, or kernel crash dumps is preferred.
- Familiarity with large-scale compute or bare-metal platforms is preferred.
- Open-source infrastructure or system software contributions are preferred.
- Experience debugging hardware and driver-level issues involving GPUs, NVLink, or InfiniBand is preferred.
Benefits
- Competitive salary and comprehensive benefits package.
- Professional growth opportunities within Nebius.
- Flexible working arrangements.
- Dynamic and collaborative work environment.
Tech Stack
Categories
Site Reliability
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.