2 hours ago
Hyderābād, IndiaSenior
Responsibilities
- Own end-to-end platform workstreams, including lifecycle upgrades across multi-region Kubernetes clusters, node operating systems, service meshes, ingress, and platform add-ons.
- Design and improve cluster provisioning and management through GitOps, lifecycle automation, and standardized Helm charts.
- Serve as a senior escalation point during on-call incidents, leading diagnosis to root cause and driving follow-up actions.
- Partner with application engineering teams on resource optimization, rightsizing, workload resiliency, and platform-blocker remediation.
- Improve container security and networking through runtime policy enforcement, secrets management, image supply-chain controls, and ingress or API gateway configuration.
- Shape the Kubernetes roadmap, estimate and sequence work, identify risks, mentor engineers, and review designs and changes.
- Create runbooks, standards, and internal documentation while converting recurring operational work into automation.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Business Administration, or equivalent experience.
- At least five years of hands-on experience operating Kubernetes in production, including at least two years across multiple clusters or regions.
- Experience with Kubernetes distributions or platforms such as Rancher/RKE2, Tanzu, OpenShift, PCF, EKS, AKS, or upstream Kubernetes.
- Deep knowledge of Kubernetes ecosystem tooling, container runtimes, cluster upgrades, node lifecycle mechanics, and platform operations.
- Proficiency in Bash and Python, or PowerShell or Go, with experience using code, Ansible, or GitOps to eliminate manual work.
- Ability to independently lead complex cross-team technical projects and make sound design decisions with minimal supervision.
- Experience mentoring or technically guiding less experienced engineers.
- Preferred experience includes container networking and security, Cluster API, infrastructure-as-code, vSphere, Kubernetes storage, monitoring and logging at scale, CI/CD pipelines, backup and disaster recovery, VMware, and public cloud platforms.
Benefits
- The role is part of a globally distributed team and requires flexibility to work across time zones.
- The position includes participation in an on-call rotation and willingness to work flexible hours.
