3 months ago
Responsibilities
- Design, implement, deploy, operate, and optimize monitoring, logging, alerting, and telemetry platforms for datacenter infrastructure.
- Build dashboards, alerts, and reports that provide visibility into system health, performance, and capacity trends.
- Apply SRE practices and develop automation for infrastructure provisioning, monitoring, and system management.
- Lead root cause analyses and post-incident reviews and drive corrective actions to improve system resilience.
- Analyze and optimize system and application performance, efficiency, and resource utilization across datacenter infrastructure.
- Collaborate with engineering teams and hardware and software vendors to define requirements, evaluate technologies, and integrate solutions.
- Implement security controls and ensure observability and reliability solutions meet organizational policies and industry standards.
- Provide hands-on troubleshooting and support for complex hardware and software issues.
- Maintain troubleshooting guides, operational documentation, and best practices.
- Continuously improve the scalability, reliability, and operational efficiency of datacenter services.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- 12+ years of progressive software engineering experience emphasizing distributed systems, cloud-native architectures, or platform operations.
- Experience managing and optimizing large-scale datacenter environments.
- Strong proficiency in Go or Python, with deep understanding of networked systems and performance optimization.
- Expert-level knowledge of Kubernetes internals, including scheduling and controllers, and containerization ecosystems.
- Experience with load balancing, service mesh, and request routing at scale.
- Proficiency with observability tools such as Prometheus, Grafana, and ELK Stack.
- Experience with SRE practices and tools such as Kubernetes, Docker, and Terraform.
- Familiarity with AWS, Azure, and GCP and their observability and reliability services.
- Preferred experience building infrastructure for LLM inference or large-scale training clusters.
- Preferred familiarity with inference, mixed precision, kernel tuning, or custom hardware accelerators.
- Preferred experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
- Preferred experience operating in regulated environments with strict security and compliance requirements.
Benefits
- Hybrid work model requiring at least 3 days in the office and allowing up to 2 days working from home per week.
- Opportunity to work on large-scale datacenter infrastructure within a fast-growing global e-commerce company.
- Employees eligible for employment protection may receive preferential treatment in accordance with applicable laws.
Tech Stack
About Coupang
Coupang is a technology and Fortune 150 company listed on the New York Stock Exchange (NYSE: CPNG) that provides retail, restaurant delivery, video streaming, and fintech services to customers around the world under brands that include Coupang, Eats, Play, Rocket Now, and Farfetch.