Coupang

Senior Staff Cloud Backend Engineer - Observability and Site Reliability

Coupang
Apply
3 months ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Design, implement, deploy, operate, and optimize monitoring, logging, alerting, and telemetry platforms for datacenter infrastructure.
  • Build dashboards, alerts, and reports that provide visibility into system health, performance, and capacity trends.
  • Apply SRE practices and develop automation for infrastructure provisioning, monitoring, and system management.
  • Lead root cause analyses and post-incident reviews and drive corrective actions to improve system resilience.
  • Analyze and optimize system and application performance, efficiency, and resource utilization across datacenter infrastructure.
  • Collaborate with engineering teams and hardware and software vendors to define requirements, evaluate technologies, and integrate solutions.
  • Implement security controls and ensure observability and reliability solutions meet organizational policies and industry standards.
  • Provide hands-on troubleshooting and support for complex hardware and software issues.
  • Maintain troubleshooting guides, operational documentation, and best practices.
  • Continuously improve the scalability, reliability, and operational efficiency of datacenter services.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
  • 12+ years of progressive software engineering experience emphasizing distributed systems, cloud-native architectures, or platform operations.
  • Experience managing and optimizing large-scale datacenter environments.
  • Strong proficiency in Go or Python, with deep understanding of networked systems and performance optimization.
  • Expert-level knowledge of Kubernetes internals, including scheduling and controllers, and containerization ecosystems.
  • Experience with load balancing, service mesh, and request routing at scale.
  • Proficiency with observability tools such as Prometheus, Grafana, and ELK Stack.
  • Experience with SRE practices and tools such as Kubernetes, Docker, and Terraform.
  • Familiarity with AWS, Azure, and GCP and their observability and reliability services.
  • Preferred experience building infrastructure for LLM inference or large-scale training clusters.
  • Preferred familiarity with inference, mixed precision, kernel tuning, or custom hardware accelerators.
  • Preferred experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
  • Preferred experience operating in regulated environments with strict security and compliance requirements.

Benefits

  • Hybrid work model requiring at least 3 days in the office and allowing up to 2 days working from home per week.
  • Opportunity to work on large-scale datacenter infrastructure within a fast-growing global e-commerce company.
  • Employees eligible for employment protection may receive preferential treatment in accordance with applicable laws.
Coupang

About Coupang

10,000+ employees

Coupang is a technology and Fortune 150 company listed on the New York Stock Exchange (NYSE: CPNG) that provides retail, restaurant delivery, video streaming, and fintech services to customers around the world under brands that include Coupang, Eats, Play, Rocket Now, and Farfetch.