28 days ago
San Diego, CA, USAStaff+
Base Salary
$108k - $180k/yr
Responsibilities
- Operate mission-critical production systems 24/7/365 and participate in on-call rotations and incident response.
- Triage and resolve production incidents, use AI-assisted log analysis and anomaly detection, reduce MTTR, and prevent recurrence.
- Manage capacity planning and resource utilization to support safe, cost-effective scaling.
- Own and operate distributed infrastructure including APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
- Design and maintain observability solutions for metrics, logs, traces, alerting, anomaly detection, and intelligent alert correlation.
- Automate operational workflows and build internal tooling using scripting and AI-assisted development tools such as Claude Code.
- Create runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Improve infrastructure reliability and performance with global engineering teams.
- Mentor Senior and mid-level SREs and lead platform modernization efforts.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- At least 6 years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Experience applying AI/LLM-powered tools to reliability engineering and building automation or internal tools with AI-assisted development tools such as Claude Code.
- Strong Linux, networking, and distributed systems fundamentals with end-to-end production debugging ability.
- Hands-on incident response, troubleshooting, and distributed-systems performance optimization experience.
- Strong software engineering skills with Python or Go and experience building automation, tooling, or platforms.
- Experience operating or supporting infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, and Zookeeper.
- Experience with Prometheus, Grafana, Zabbix, observability, monitoring, and performance analysis.
- Familiarity with Git, configuration management tools such as Ansible, and CI/CD pipelines.
- Strong ownership, systematic problem-solving, communication, and collaboration skills.
- Preferred: bilingual Mandarin and English fluency; Kubernetes Administrator certification or equivalent experience; experience with Hadoop, Yarn, HBase, Hive, or Spark.
Benefits
- Bonus and RSU eligibility; healthcare, insurance, spending accounts, Employee Assistance Program, business travel accident insurance, 401(k) with discretionary company match, paid vacation and holidays, employee discounts, catered lunches, select-location gym access and dog-friendly offices, company events, and complimentary snacks and beverages.
Tech Stack
AnsibleApache HadoopApache HBaseApache HiveApache KafkaApache SparkConsulElasticsearchGitGoGrafanaKubernetesLinuxPrometheusPythonRedisYarn
Categories
Site Reliability
