SHEIN

Staff Site Reliability Engineer

SHEIN
Apply
28 days ago
San Diego, CA, USAStaff+

Base Salary

$108k - $180k/yr

Responsibilities

  • Operate mission-critical production systems 24/7/365 and participate in on-call rotations and incident response.
  • Triage and resolve production incidents, use AI-assisted log analysis and anomaly detection, reduce MTTR, and prevent recurrence.
  • Manage capacity planning and resource utilization to support safe, cost-effective scaling.
  • Own and operate distributed infrastructure including APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
  • Design and maintain observability solutions for metrics, logs, traces, alerting, anomaly detection, and intelligent alert correlation.
  • Automate operational workflows and build internal tooling using scripting and AI-assisted development tools such as Claude Code.
  • Create runbooks, architecture diagrams, operational procedures, and on-call playbooks.
  • Improve infrastructure reliability and performance with global engineering teams.
  • Mentor Senior and mid-level SREs and lead platform modernization efforts.

Requirements

  • Bachelor’s degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
  • At least 6 years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
  • Experience applying AI/LLM-powered tools to reliability engineering and building automation or internal tools with AI-assisted development tools such as Claude Code.
  • Strong Linux, networking, and distributed systems fundamentals with end-to-end production debugging ability.
  • Hands-on incident response, troubleshooting, and distributed-systems performance optimization experience.
  • Strong software engineering skills with Python or Go and experience building automation, tooling, or platforms.
  • Experience operating or supporting infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
  • Experience with Prometheus, Grafana, Zabbix, observability, monitoring, and performance analysis.
  • Familiarity with Git, configuration management tools such as Ansible, and CI/CD pipelines.
  • Strong ownership, systematic problem-solving, communication, and collaboration skills.
  • Preferred: bilingual Mandarin and English fluency; Kubernetes Administrator certification or equivalent experience; experience with Hadoop, Yarn, HBase, Hive, or Spark.

Benefits

  • Bonus and RSU eligibility; healthcare, insurance, spending accounts, Employee Assistance Program, business travel accident insurance, 401(k) with discretionary company match, paid vacation and holidays, employee discounts, catered lunches, select-location gym access and dog-friendly offices, company events, and complimentary snacks and beverages.

Tech Stack

AnsibleApache HadoopApache HBaseApache HiveApache KafkaApache SparkConsulElasticsearchGitGoGrafanaKubernetesLinuxPrometheusPythonRedisYarn

Categories

Site Reliability
SHEIN

About SHEIN

10,000+ employees
Contact me