FriendliAI

Software Engineer – Cloud Infrastructure

FriendliAI
Apply
24 days ago

Responsibilities

  • Own the architecture, topology, lifecycle, and zero-downtime upgrades of a multi-cluster, multi-tenant Kubernetes fleet.
  • Extend Kubernetes with custom controllers, operators, and CRDs.
  • Design GPU scheduling, capacity strategy, quotas, autoscaling, scale-to-zero, and cold-start reduction for inference workloads.
  • Own the Kubernetes network data plane, including CNI, IPAM, DNS, ingress, load balancing, and cross-AZ, cross-region, and cross-cluster connectivity.
  • Operate service-mesh routing, mTLS, and traffic policies, and debug production networking issues such as packet loss, conntrack exhaustion, MTU mismatches, and DNS latency.
  • Define platform SLOs, lead post-incident hardening, and deliver infrastructure as code with Terraform, Helm, and GitOps.
  • Partner with inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.

Requirements

  • 5+ years of experience designing, building, and operating large-scale Kubernetes infrastructure in production.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • Proven experience operating large-scale, high-traffic network services in production.
  • Deep understanding of Kubernetes internals, including the API server, scheduler, controller loops, kubelet, and etcd.
  • Strong command of Kubernetes and cloud networking, including CNI, kube-proxy or eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
  • Proficiency with AWS, Terraform, Helm, and Ansible.
  • Programming ability in Go or Python for infrastructure tooling and automation.
  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.
  • Clear written and verbal communication and the ability to document architectural decisions.
  • Preferred experience includes large-scale Kubernetes operations, Cilium and eBPF, Kubespray, NVIDIA GPU Operator, device plugins, Dynamic Resource Allocation, RDMA/RoCE, InfiniBand, EFA, SR-IOV, NCCL tuning, multi-cloud or bare-metal Kubernetes, and contributions to Kubernetes, Cilium, Istio, or other CNCF projects.

Benefits

  • Flexible working hours.
  • Daily lunch and dinner, unlimited snacks and beverages.
  • Supportive and highly collaborative work environment.
  • Health check-up support and top-tier equipment and hardware support.
  • Competitive compensation, startup equity, health insurance, and other benefits.
  • Opportunity to work on generative AI infrastructure at scale.

Tech Stack

Categories

FriendliAI

About FriendliAI

51-200 employees
Contact me