4 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect and build an AI Inference Gateway platform for large-scale machine learning and generative AI workloads.
- Develop request routing, model endpoint abstraction, traffic shaping, load balancing, failover, caching, and policy enforcement mechanisms.
- Drive the technical vision and roadmap for secure, reliable AI inference infrastructure across cloud and on-premises environments.
- Build infrastructure components in Go, Java, or Python with emphasis on performance, resiliency, and operational excellence.
- Design multi-tenant capabilities for authentication, authorization, quota management, cost attribution, rate limiting, and governance.
- Partner with ML Platform, Model Serving, Data, and Product Engineering teams to deploy and operate AI workloads.
- Lead architecture and design reviews, raise engineering standards, and mentor senior engineers across multiple teams.
- Optimize latency, throughput, and infrastructure efficiency for large-scale inference workloads.
- Define observability standards using metrics, tracing, logging, and SLO-based operations.
- Investigate production issues, conduct root-cause analysis, and implement durable architectural solutions.
- Collaborate with engineering leaders to establish common platform standards and enable AI innovation.
Requirements
- 8+ years of professional software development experience.
- 5+ years designing and operating large-scale distributed systems in production.
- Strong programming expertise in one or more of Go, Java, or Python.
- Experience building highly available, mission-critical platform or infrastructure services.
- Experience designing API platforms, service gateways, service meshes, or large-scale networking infrastructure.
- Deep understanding of distributed systems design and cloud-native technologies.
- Production experience with Kubernetes and containerized workloads.
- Experience operating services on AWS, Azure, or GCP.
- Strong understanding of observability, reliability engineering, capacity planning, and production operations.
- Preferred experience building AI/ML inference platforms, LLM gateways, model-serving infrastructure, or GPU-accelerated workloads.
- Preferred expertise with Gateway API, ingress controllers, Istio, Linkerd, Envoy, and platform networking.
- Preferred experience with vLLM, Triton Inference Server, TensorRT-LLM, Ray Serve, KServe, SGLang, or similar inference-serving technologies.
- Preferred understanding of high-performance networking, gRPC, HTTP/2, streaming protocols, API gateways, and service proxy architectures.
- Preferred experience with large-scale traffic management, latency and throughput optimization, distributed data systems, and asynchronous programming.
Benefits
- Hybrid, onsite, or remote work models are available depending on role requirements.
- The hybrid model requires at least 3 days in the office per week and allows up to 2 days working from home, depending on role requirements.
- Some businesses may require additional time in the office.
Tech Stack
AmbassadorApache CassandraApache KafkaArgo CDAWSAzureGoGoogle Cloud PlatformgRPCIstioJavaKubernetesMongoDBPythonRedis
Categories
About Coupang
Coupang builds and operates a South Korea–focused e-commerce marketplace with an end-to-end logistics network (Rocket Delivery), plus food delivery, video streaming, and fintech under brands such as Coupang, Eats, and Play. Revenue comes from first-party retail, third-party marketplace services, advertising, and memberships (Rocket WOW). Founded in 2010, the company is headquartered in Seattle and is publicly listed on the NYSE (CPNG).
