Responsibilities
- Own datacenter compute and storage lifecycle, including bare-metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.
- Build automation for PXE/firmware workflows, dynamic inventory, image generation, configuration drift detection, and automated recovery.
- Define and deliver reliability targets for provisioning, node commissioning, cluster upgrades, and mean time to recover.
- Lead infrastructure incident response, on-call rotations, incident command, postmortems, and follow-up actions.
- Maintain monitoring, alerting, and dashboards for hardware, hypervisors, storage, Kubernetes control planes, and autoscaling.
- Manage infrastructure cost, capacity, reclamation, firmware and hypervisor patching, cold standby, and failover procedures.
- Perform datacenter racking, cabling, hardware troubleshooting, forensic log capture, and coordination of physical repairs.
Requirements
- At least 5 years of engineering experience, including at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
- Hands-on expertise with bare-metal provisioning and imaging using PXE/iPXE, IPMI, and Redfish.
- Experience operating KVM/QEMU or ESXi hypervisors and Ceph, NVMeoF, or SAN storage systems at scale.
- Production Kubernetes operations experience covering provisioning, upgrades, highly available control planes, cluster APIs, CNI and CSI troubleshooting, and multi-cluster workload scheduling.
- Production automation and coding skills in Python, Go, or Rust, plus experience with Terraform, Ansible, Helm, and CI/CD pipelines.
- Strong networking fundamentals including VLANs, BGP/EVPN leaf-spine networks, LACP, routing, cluster networking, and storage-fabric troubleshooting.
- Experience with on-call operations, postmortems, SLIs/SLOs, and reliability improvements.
- Ability to work hands-on in datacenters and operate in regulated or safety-sensitive environments with change-control and audit processes.
- Ability to work regularly onsite in South San Francisco, travel occasionally to datacenter or field locations, and collaborate across technical and operations teams.
Benefits
- South San Francisco-based role with regular onsite presence and in-office cadence.
- Participation in on-call rotations and occasional travel to partner datacenters or field sites.
- Equal opportunity workplace committed to diversity and inclusion.
About Zipline
Zipline was founded to create the first logistics system that serves all humans equally. Our aim is to solve the world’s most urgent and complex access challenges. Leveraging expertise in robotics and autonomy, Zipline designs, manufactures and operates the world’s largest automated delivery system. Zipline serves tens of millions of people around the world and is making good on the promise of building an equitable and more resilient global supply chain. From powering Rwanda’s national blood delivery network and Ghana’s COVID-19 vaccine distribution, to providing on-demand home delivery for Walmart and enabling leading healthcare providers to bring care into the home in the United States, Zipline is transforming the way goods move. By transitioning to clean, electric, instant logistics, we can decarbonize delivery, decrease road congestion, and reduce fossil fuel consumption and air pollution, while providing equitable access for billions of people. The technology is complex but the idea is simple: a teleportation service that delivers what you need, when you need it. Zipline is inspiring people, governments, and businesses to imagine what is possible when goods can move as seamlessly as information. To join the team, check out our career page: https://flyzipline.com/careers/
