Nebius

Senior Site Reliability Engineer — AI Studio (Inference Platform)

Nebius
Apply
9 months ago
Remote, United States +5 moreSenior

Responsibilities

  • Own the reliability, performance, and observability of the entire inference stack.
  • Design and refine telemetry pipelines for metrics, logs, and traces at large scale.
  • Tune Kubernetes autoscalers to improve GPU efficiency.
  • Create Terraform modules that embed resilience into new clusters.
  • Harden request-routing and retry logic to handle transient failures.
  • Build automation and runbooks to detect, isolate, and remediate incidents quickly.
  • Drive post-mortems and improvements that prevent recurring incidents.
  • Scale the platform while meeting aggressive cost and reliability targets.

Requirements

  • Deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and infrastructure-as-code.
  • Comfortable scripting in Python or Bash.
  • Understanding of alert design and service-level objectives for high-throughput APIs.
  • Production experience with distributed back-end systems and their failure modes.
  • Experience with GPU-heavy workloads using vLLM, Triton, Ray, or another accelerator stack is valuable.
  • Background in MLOps or model-hosting platforms is valuable.
  • Ability to debug performance from the kernel through the application layer and collaborate with software engineers.

Benefits

  • Competitive salary and comprehensive benefits package.
  • Opportunities for professional growth within Nebius.
  • Flexible working arrangements.
  • Dynamic and collaborative work environment that values initiative and innovation.

Tech Stack

BashGrafanaKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Nebius

About Nebius

1,001-5,000 employees

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Contact me