1 hour ago
Bengaluru, IndiaStaff+
Responsibilities
- Provision and harden bare-metal Linux servers across multiple metros using PXE/iPXE and cloud-init automation.
- Bring up and operate Kubernetes clusters, including control planes, etcd, worker nodes, CNI configuration, validation, node lifecycle, certificates, capacity, and cluster alerts.
- Maintain GitOps-based cluster and fleet configuration, including registration, RBAC, policies, and production/nonproduction isolation.
- Extend deployment-pipeline validation and test coverage for the software-defined data plane.
- Execute rolling OS and Kubernetes upgrades, patch fleets, track CVEs, take etcd snapshots, and conduct restore drills.
- Respond to control-plane, node, and cluster incidents through a follow-the-sun on-call rotation.
- Build dashboards and alerts, reduce alert noise, automate recurring toil, write incident reports, contribute to postmortems, and maintain runbooks.
- Collaborate with hardware, networking, and data-plane teams to isolate faults involving servers, NICs, disks, firmware, and network layers.
Requirements
- Hands-on Linux/Ubuntu administration, bare-metal provisioning, PXE/iPXE or equivalent imaging, cloud-init, and OS hardening experience.
- Working knowledge of BMC/IPMI, SMART data, platform sensors, hardware fault isolation, and vendor replacement processes.
- Production experience operating bare-metal Kubernetes clusters, including etcd, node lifecycle, CNI configuration, and cluster troubleshooting; RKE2 or kubeadm experience is preferred.
- Day-to-day experience with Ansible, Terraform, ArgoCD, Fleet, or comparable infrastructure-as-code and GitOps tooling.
- Working knowledge of VLANs, BGP, bonded NICs, and SR-IOV sufficient to interpret configurations and isolate network-layer problems.
- Experience executing production OS and Kubernetes upgrades from established runbooks.
- Experience in an on-call operations role, responding to production incidents and documenting them afterward.
- Ability and inclination to automate repeated manual work and improve operational runbooks.
- Preferred experience with Rancher, Rancher Prime, or comparable multi-cluster management platforms.
- Preferred familiarity with DPDK-based data planes such as 6WIND VSR, SR-IOV NIC tuning, SDN, private-connectivity platforms, or Equinix Fabric and VyOS.
- Preferred familiarity with LLM provider APIs and AI/agent gateway proxying patterns.
- Preferred experience operating infrastructure across geographically distributed sites.
Tech Stack
Categories
About Equinix
Equinix (Nasdaq: EQIX) is the world’s digital infrastructure company®, enabling digital leaders to harness a trusted platform to bring together and interconnect the foundational infrastructure that powers their success. Equinix enables today’s businesses to access all the right places, partners and possibilities they need to accelerate advantage. With Equinix, they can scale with agility, speed the launch of digital services, deliver world-class experiences and multiply their value.
