Software Engineer - Cloud Infrastructure
FriendliAI29 days ago
Seoul, Korea, SouthSenior
Responsibilities
- Own the architecture, topology, lifecycle, and zero-downtime upgrades of a multi-cluster, multi-tenant Kubernetes fleet.
- Extend Kubernetes with custom controllers, operators, and CRDs.
- Design GPU scheduling, capacity management, topology-aware placement, node pools, priority, preemption, and tenant quotas.
- Build queue-driven pod autoscaling, node autoscaling, scale-to-zero, and cold-start reduction systems.
- Own the Kubernetes network data plane, including CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
- Design cross-AZ, cross-region, and cross-cluster connectivity and operate the service mesh for routing, mTLS, and traffic policy.
- Debug production network issues and implement permanent fixes.
- Define platform SLOs, lead post-incident hardening, and deliver infrastructure as code with Terraform, Helm, and GitOps.
- Partner with inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.
Requirements
- 5+ years of experience designing, building, and operating large-scale Kubernetes infrastructure in production.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- Experience operating large-scale, high-traffic network services in production.
- Deep understanding of Kubernetes internals, including the API server, scheduler, controller loops, kubelet, and etcd.
- Strong command of Kubernetes and cloud networking, including CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
- Proficiency with AWS, Terraform, Helm, and Ansible.
- Programming skills in Go or Python for infrastructure tooling and automation.
- Strong debugging skills across distributed systems, containers, and the Linux networking stack.
- Strong written and verbal communication skills, including documenting architectural decisions.
- Preferred experience with large-scale Kubernetes in high-traffic domains, Cilium, eBPF, Kubespray, GPU orchestration, NVIDIA GPU Operator, device plugins, DRA, RDMA/RoCE, InfiniBand, EFA, SR-IOV, NCCL tuning, multi-cloud, hybrid-cloud, bare-metal Kubernetes, or contributions to Kubernetes, Cilium, Istio, or other CNCF projects.
Benefits
- Flexible working hours.
- Daily lunch and dinner, unlimited snacks and beverages.
- Supportive and highly collaborative work environment.
- Health check-up support and top-tier equipment/hardware support.
- Health insurance, startup equity, and other benefits.
- The posting describes a small, fast-moving team working on generative AI infrastructure.