Nebius

Site Reliability Engineer in Network Infrastructure

Nebius
Apply
2 months ago
Remote, EMEA or Amsterdam, NetherlandsSenior

Responsibilities

  • Define and own reliability goals, SLIs, SLOs, availability targets, and error budgets for network services and critical paths.
  • Drive reliability improvements across network services, site readiness, inter-site connectivity, and operational standards.
  • Own incident response, lead investigations and postmortems, and implement durable fixes.
  • Build and improve metrics, logs, traces, alerting, and debugging workflows.
  • Design safer network change workflows using automation, testing and staging environments, canarying, rollbacks, and auditability.
  • Collaborate with network engineers and platform teams to embed operability into system designs and keep operations practical.

Requirements

  • Strong production Linux fundamentals and a structured approach to debugging complex systems.
  • Solid understanding of networking fundamentals, including control and data planes, latency, packet loss, and failure domains.
  • Hands-on experience operating and continuously improving high-availability systems.
  • Ability to write and maintain software and automation, particularly in Go or Python.
  • Experience with modern infrastructure tooling, infrastructure as code, CI/CD, and container platforms.
  • Preferred experience with high-throughput traffic processing systems such as load balancers, tunneling or decapsulation, and NAT64.
  • Preferred background in low-level networking performance and debugging, including eBPF/XDP, DPDK, perf/ftrace, or kernel networking internals.
  • Preferred experience building network-safe delivery pipelines and large-scale network observability and telemetry systems.

Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

Tech Stack

Categories

DevOpsSite Reliability
Nebius

About Nebius

1,001-5,000 employees

Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.

Contact me