1 hour ago
Bengaluru, IndiaStaff+
Responsibilities
- Provision and harden Debian/Ubuntu operating systems on bare-metal servers across multiple metro locations using PXE/iPXE and cloud-init.
- Build and maintain zero-touch infrastructure provisioning pipelines and collaborate on BIOS, firmware, NIC, disk, and hardware fault diagnosis.
- Deploy and operate multi-cluster Kubernetes environments on bare metal, including highly available kube-apiserver instances, worker nodes, CNI configuration, and cluster validation.
- Maintain Layer 4 load balancing or keepalived VIPs, manage etcd snapshots and restores, and execute rolling OS and Kubernetes upgrades without unplanned downtime.
- Build and own GitOps and CI/CD pipelines for software-defined data-plane deployment, including throughput and latency validation gates.
- Administer centralized multi-cluster management planes, including RBAC, cluster registration, policy enforcement, and production/non-production isolation.
- Operate and troubleshoot networks involving VLANs, BGP, bonded NICs, SR-IOV, private connectivity, firewall rules, and segmentation.
- Own observability, alerting thresholds, operational runbooks, on-call responsibilities, infrastructure incident response, and blameless postmortems.
- Automate recurring operational tasks and reduce failure modes through improved tooling, runbooks, and alerting.
Requirements
- Deep hands-on experience provisioning and hardening Linux/Ubuntu systems at scale with PXE/iPXE, cloud-init, BIOS/firmware management, and bare-metal infrastructure.
- Experience with BMC/IPMI, SMART data, platform sensors, and diagnosing disk, memory, NIC, and power faults while coordinating vendor replacements.
- Strong experience deploying and operating RKE2 or kubeadm Kubernetes clusters on bare metal, including etcd, kube-apiserver high availability, CNI, Multus, SR-IOV, and multi-cluster fleet management.
- Production experience managing multiple clusters with Rancher Prime, Rancher, or an equivalent management plane.
- Experience with GitOps or pipeline-driven infrastructure deployment and tools such as Ansible, Terraform, ArgoCD, or Fleet.
- Hands-on networking experience with VLANs, BGP, bonded NICs, SR-IOV, and private connectivity, including troubleshooting across infrastructure boundaries.
- Demonstrated experience performing zero-downtime rolling upgrades of operating systems and Kubernetes across multi-node fleets.
- Experience operating highly available platforms, leading infrastructure incident response, participating in on-call rotations, and using automation-first, auditable operational practices.
- Preferred familiarity with 6WIND VSR or other DPDK-based data planes, SR-IOV NIC tuning, and data-plane performance validation.
- Preferred familiarity with SDN or private-connectivity platforms such as Equinix Fabric or VyOS.
- Preferred familiarity with LLM provider APIs and AI/agent gateway proxying patterns, plus experience operating infrastructure across geographically distributed sites.
Tech Stack
Categories
About Equinix
Equinix (Nasdaq: EQIX) is the world’s digital infrastructure company®, enabling digital leaders to harness a trusted platform to bring together and interconnect the foundational infrastructure that powers their success. Equinix enables today’s businesses to access all the right places, partners and possibilities they need to accelerate advantage. With Equinix, they can scale with agility, speed the launch of digital services, deliver world-class experiences and multiply their value.
