over 3 years ago
Bengaluru, IndiaSenior
Responsibilities
- Reduce Kubernetes overprovisioning, drive right-sizing programs, and maintain infrastructure cost telemetry.
- Run GPU throughput optimization experiments on on-premise GPU clusters in partnership with AI service owners.
- Build tooling, dashboards, and processes that enable backend teams to manage their own cost and reliability budgets.
- Ensure new and offline system flows are properly instrumented for cost-at-scale and reliability.
- Contribute to defined platform-security workstreams and platform-security changes.
- Operate across backend engineering, infrastructure operations, and FinOps in hands-on hybrid environments.
Requirements
- 4–5 years of hands-on systems experience.
- Production experience with Python and Go or Rust, including end-to-end service ownership and the ability to reason about backend code across teams.
- Kubernetes-at-scale expertise, including scheduler behavior, resource requests and limits, HPA, VPA, node pool design, and cost-aware autoscaling with Cast AI, Karpenter, or equivalent.
- Fluency with GCP and Terraform, plus experience with CI/CD and hybrid environments containing on-premise GPU clusters.
- Understanding of throughput profiling, batching, KV-cache behavior, inference-server tuning, and GPU-utilization metrics.
- Experience with metrics, traces, logs, SLOs, and disciplined reliability instrumentation.
- Demonstrated ability to translate infrastructure choices into measurable cost outcomes.
- Ability to take on platform-security workstreams without constant handoff to a DevOps team.
Categories
DevOpsSite Reliability
