8 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect and build Coupang's AI Inference Gateway platform for large-scale machine learning and generative AI workloads.
- Develop request routing, model endpoint abstraction, traffic shaping, load balancing, failover, caching, and policy enforcement mechanisms.
- Design scalable, secure, reliable, and multi-tenant inference infrastructure across cloud and on-premises environments.
- Implement authentication, authorization, quota management, cost attribution, rate limiting, and governance controls.
- Partner with ML Platform, Model Serving, Data, and Product Engineering teams to deploy and operate AI workloads.
- Optimize performance, latency, throughput, and infrastructure efficiency for CPU and GPU inference workloads.
- Define observability standards using metrics, tracing, logging, and SLO-based operations.
- Investigate production issues, perform root-cause analysis, and implement long-term architectural solutions.
- Lead architecture and design reviews, establish platform standards, and mentor senior engineers across teams.
Requirements
- At least 8 years of professional software development experience.
- At least 5 years of experience designing and operating large-scale distributed systems in production.
- Strong programming expertise in Go, Java, or Python.
- Experience building highly available, mission-critical platform or infrastructure services.
- Experience with API platforms, service gateways, service mesh, or large-scale networking infrastructure.
- Deep understanding of microservices, distributed systems design, cloud-native technologies, observability, reliability engineering, capacity planning, and production operations.
- Production experience with Kubernetes and containerized workloads.
- Experience operating services on AWS, Azure, or GCP.
- Preferred experience with AI/ML inference platforms, LLM gateways, model serving infrastructure, or GPU-accelerated workloads.
- Preferred expertise with Gateway API, Ingress Controllers, Istio, Linkerd, Envoy, and platform networking.
- Preferred experience with vLLM, Triton Inference Server, TensorRT-LLM, Ray Serve, KServe, SGLang, or similar inference-serving technologies.
- Preferred experience with high-performance networking, gRPC, HTTP/2, streaming protocols, API gateways, service proxies, traffic management, and distributed data systems.
- Preferred familiarity with GenAI, LLM ecosystems, model deployment, prompt routing, RAG systems, model observability, and AI governance.
Benefits
- Hybrid, onsite, or remote work options are available depending on role requirements.
- Coupang's hybrid model generally requires at least 3 days per week in the office and allows up to 2 days working from home.
- Some businesses may require additional in-office time based on the nature of the work.
Tech Stack
AmbassadorApache CassandraApache KafkaArgo CDAWSAzureGoGoogle Cloud PlatformgRPCIstioJavaKubernetesMongoDBPythonRedis
Categories
About Coupang
Coupang builds and operates a South Korea–focused e-commerce marketplace with an end-to-end logistics network (Rocket Delivery), plus food delivery, video streaming, and fintech under brands such as Coupang, Eats, and Play. Revenue comes from first-party retail, third-party marketplace services, advertising, and memberships (Rocket WOW). Founded in 2010, the company is headquartered in Seattle and is publicly listed on the NYSE (CPNG).
