Level AI

Senior Site Reliability Engineer (Noida, BLR, India)

Level AI
Apply
over 3 years ago
Bengaluru, IndiaSenior

Responsibilities

  • Reduce Kubernetes overprovisioning, drive right-sizing programs, and maintain infrastructure cost telemetry.
  • Run GPU throughput optimization experiments on on-premise GPU clusters in partnership with AI service owners.
  • Build tooling, dashboards, and processes that enable backend teams to manage their own cost and reliability budgets.
  • Ensure new and offline system flows are properly instrumented for cost-at-scale and reliability.
  • Contribute to defined platform-security workstreams and platform-security changes.
  • Operate across backend engineering, infrastructure operations, and FinOps in hands-on hybrid environments.

Requirements

  • 4–5 years of hands-on systems experience.
  • Production experience with Python and Go or Rust, including end-to-end service ownership and the ability to reason about backend code across teams.
  • Kubernetes-at-scale expertise, including scheduler behavior, resource requests and limits, HPA, VPA, node pool design, and cost-aware autoscaling with Cast AI, Karpenter, or equivalent.
  • Fluency with GCP and Terraform, plus experience with CI/CD and hybrid environments containing on-premise GPU clusters.
  • Understanding of throughput profiling, batching, KV-cache behavior, inference-server tuning, and GPU-utilization metrics.
  • Experience with metrics, traces, logs, SLOs, and disciplined reliability instrumentation.
  • Demonstrated ability to translate infrastructure choices into measurable cost outcomes.
  • Ability to take on platform-security workstreams without constant handoff to a DevOps team.
Level AI

About Level AI

201-500 employees
Contact me