about 3 hours ago
Remote, Worldwide or Amsterdam, NetherlandsMid Level / Senior
Responsibilities
- Define and own reliability goals for network services and critical paths.
- Drive reliability improvements across the network, including site readiness and operational standards.
- Own incident response, lead investigations, and implement durable fixes.
- Build and evolve observability tools for metrics, logs, and alerting.
- Design safer change workflows with automation and CI/CD processes.
- Collaborate with network engineers to embed operability into designs.
Requirements
- Strong production Linux fundamentals and structured debugging skills.
- Solid understanding of networking basics and failure modes.
- Hands-on experience with high-availability systems and their improvement.
- Ability to write and maintain software/automation, preferably in Go or Python.
- Experience with modern infrastructure tooling and automating operational workflows.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams