FriendliAI

Software Engineer - Cloud Infrastructure

FriendliAI
Apply
29 days ago
Seoul, Korea, SouthSenior

Responsibilities

  • Own the architecture, topology, lifecycle, and zero-downtime upgrades of a multi-cluster, multi-tenant Kubernetes fleet.
  • Extend Kubernetes with custom controllers, operators, and CRDs.
  • Design GPU scheduling, capacity management, topology-aware placement, node pools, priority, preemption, and tenant quotas.
  • Build queue-driven pod autoscaling, node autoscaling, scale-to-zero, and cold-start reduction systems.
  • Own the Kubernetes network data plane, including CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ, cross-region, and cross-cluster connectivity and operate the service mesh for routing, mTLS, and traffic policy.
  • Debug production network issues and implement permanent fixes.
  • Define platform SLOs, lead post-incident hardening, and deliver infrastructure as code with Terraform, Helm, and GitOps.
  • Partner with inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.

Requirements

  • 5+ years of experience designing, building, and operating large-scale Kubernetes infrastructure in production.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • Experience operating large-scale, high-traffic network services in production.
  • Deep understanding of Kubernetes internals, including the API server, scheduler, controller loops, kubelet, and etcd.
  • Strong command of Kubernetes and cloud networking, including CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
  • Proficiency with AWS, Terraform, Helm, and Ansible.
  • Programming skills in Go or Python for infrastructure tooling and automation.
  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.
  • Strong written and verbal communication skills, including documenting architectural decisions.
  • Preferred experience with large-scale Kubernetes in high-traffic domains, Cilium, eBPF, Kubespray, GPU orchestration, NVIDIA GPU Operator, device plugins, DRA, RDMA/RoCE, InfiniBand, EFA, SR-IOV, NCCL tuning, multi-cloud, hybrid-cloud, bare-metal Kubernetes, or contributions to Kubernetes, Cilium, Istio, or other CNCF projects.

Benefits

  • Flexible working hours.
  • Daily lunch and dinner, unlimited snacks and beverages.
  • Supportive and highly collaborative work environment.
  • Health check-up support and top-tier equipment/hardware support.
  • Health insurance, startup equity, and other benefits.
  • The posting describes a small, fast-moving team working on generative AI infrastructure.

Tech Stack

Categories

FriendliAI

About FriendliAI

51-200 employees
Contact me