2 months ago
Remote, EMEA or Amsterdam, NetherlandsSenior
Responsibilities
- Define and own reliability goals, SLIs, SLOs, availability targets, and error budgets for network services and critical paths.
- Drive reliability improvements across network services, site readiness, inter-site connectivity, and operational standards.
- Own incident response, lead investigations and postmortems, and implement durable fixes.
- Build and improve metrics, logs, traces, alerting, and debugging workflows.
- Design safer network change workflows using automation, testing and staging environments, canarying, rollbacks, and auditability.
- Collaborate with network engineers and platform teams to embed operability into system designs and keep operations practical.
Requirements
- Strong production Linux fundamentals and a structured approach to debugging complex systems.
- Solid understanding of networking fundamentals, including control and data planes, latency, packet loss, and failure domains.
- Hands-on experience operating and continuously improving high-availability systems.
- Ability to write and maintain software and automation, particularly in Go or Python.
- Experience with modern infrastructure tooling, infrastructure as code, CI/CD, and container platforms.
- Preferred experience with high-throughput traffic processing systems such as load balancers, tunneling or decapsulation, and NAT64.
- Preferred background in low-level networking performance and debugging, including eBPF/XDP, DPDK, perf/ftrace, or kernel networking internals.
- Preferred experience building network-safe delivery pipelines and large-scale network observability and telemetry systems.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
