GrepJob
Nebius

Site Reliability Engineer in Network Infrastructure

Nebius
Apply
about 3 hours ago
Remote, Worldwide or Amsterdam, NetherlandsMid Level / Senior

Responsibilities

  • Define and own reliability goals for network services and critical paths.
  • Drive reliability improvements across the network, including site readiness and operational standards.
  • Own incident response, lead investigations, and implement durable fixes.
  • Build and evolve observability tools for metrics, logs, and alerting.
  • Design safer change workflows with automation and CI/CD processes.
  • Collaborate with network engineers to embed operability into designs.

Requirements

  • Strong production Linux fundamentals and structured debugging skills.
  • Solid understanding of networking basics and failure modes.
  • Hands-on experience with high-availability systems and their improvement.
  • Ability to write and maintain software/automation, preferably in Go or Python.
  • Experience with modern infrastructure tooling and automating operational workflows.

Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

Tech Stack

Categories