Software Engineer – Cloud Infrastructure
FriendliAI24 days ago
Responsibilities
- Own the architecture, topology, lifecycle, and zero-downtime upgrades of a multi-cluster, multi-tenant Kubernetes fleet.
- Extend Kubernetes with custom controllers, operators, and CRDs.
- Design GPU scheduling, capacity strategy, quotas, autoscaling, scale-to-zero, and cold-start reduction for inference workloads.
- Own the Kubernetes network data plane, including CNI, IPAM, DNS, ingress, load balancing, and cross-AZ, cross-region, and cross-cluster connectivity.
- Operate service-mesh routing, mTLS, and traffic policies, and debug production networking issues such as packet loss, conntrack exhaustion, MTU mismatches, and DNS latency.
- Define platform SLOs, lead post-incident hardening, and deliver infrastructure as code with Terraform, Helm, and GitOps.
- Partner with inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.
Requirements
- 5+ years of experience designing, building, and operating large-scale Kubernetes infrastructure in production.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- Proven experience operating large-scale, high-traffic network services in production.
- Deep understanding of Kubernetes internals, including the API server, scheduler, controller loops, kubelet, and etcd.
- Strong command of Kubernetes and cloud networking, including CNI, kube-proxy or eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
- Proficiency with AWS, Terraform, Helm, and Ansible.
- Programming ability in Go or Python for infrastructure tooling and automation.
- Strong debugging skills across distributed systems, containers, and the Linux networking stack.
- Clear written and verbal communication and the ability to document architectural decisions.
- Preferred experience includes large-scale Kubernetes operations, Cilium and eBPF, Kubespray, NVIDIA GPU Operator, device plugins, Dynamic Resource Allocation, RDMA/RoCE, InfiniBand, EFA, SR-IOV, NCCL tuning, multi-cloud or bare-metal Kubernetes, and contributions to Kubernetes, Cilium, Istio, or other CNCF projects.
Benefits
- Flexible working hours.
- Daily lunch and dinner, unlimited snacks and beverages.
- Supportive and highly collaborative work environment.
- Health check-up support and top-tier equipment and hardware support.
- Competitive compensation, startup equity, health insurance, and other benefits.
- Opportunity to work on generative AI infrastructure at scale.