28 days ago
San Diego, CA, USASenior
Base Salary
$92k - $149k/yr
Responsibilities
- Operate SHEIN’s mission-critical production systems continuously and participate in on-call rotations and incident response.
- Triage, resolve, and prevent production incidents while improving MTTR and system resilience.
- Manage capacity planning, resource utilization, scalability, and cost-effective infrastructure operations.
- Own and operate distributed infrastructure including APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
- Design and maintain observability solutions for metrics, logs, traces, alerting, anomaly detection, and alert correlation.
- Automate operational workflows and build internal tooling using scripting, software engineering, and AI-assisted development tools.
- Maintain runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Collaborate with global engineering teams to improve infrastructure reliability, performance, and operational discipline.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- At least 3 years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Strong foundations in Linux, networking, distributed systems, incident response, troubleshooting, and performance optimization.
- Experience applying AI/LLM-powered tools to reliability engineering and building automation or internal tools with AI-assisted development tools such as Claude Code.
- Strong software engineering skills using languages such as Python or Go.
- Experience operating or supporting infrastructure components including APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
- Experience with observability and monitoring systems such as Prometheus, Grafana, and Zabbix.
- Familiarity with Git, CI/CD pipelines, and configuration management tools such as Ansible.
- Strong ownership, systematic problem-solving, communication, and collaboration skills.
- Preferred qualifications include Mandarin and English fluency, Kubernetes Administrator certification or equivalent experience, and experience with Hadoop, Yarn, HBase, Hive, or Spark.
Benefits
- Bonus eligibility.
- Medical, dental, vision, and prescription drug healthcare coverage.
- Health Savings Account with employer funding and Flexible Spending Accounts.
- Company-paid basic life/AD&D insurance and short- and long-term disability coverage.
- Voluntary life/AD&D, hospital indemnity, critical illness, and accident benefits.
- Employee Assistance Program and business travel accident insurance.
- 401(k) savings plan with discretionary company match and access to a financial advisor.
- Vacation, paid holidays, floating holidays, and sick days.
- Employee discounts, free weekly catered lunch, snacks and beverages, and free swag giveaways.
- Dog-friendly offices and free gym access at select locations, plus company events and an annual holiday party.
Tech Stack
AnsibleApache HadoopApache HBaseApache HiveApache KafkaApache SparkConsulElasticsearchGitGoGrafanaKubernetesLinuxPrometheusPythonRedisYarn
Categories
Site Reliability
