ByteDance

Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance
Apply
14 hours ago
San Jose, CA, USASenior
H1B sponsor

Responsibilities

  • Respond to, troubleshoot, and resolve production incidents while serving as incident commander for critical events.
  • Lead blameless post-incident reviews and ensure corrective actions are implemented.
  • Define and maintain SLOs, SLAs, and error budgets for critical data services.
  • Drive capacity planning, performance tuning, resource management, and infrastructure cost optimization.
  • Design automation and AI-agent workflows to reduce operational toil and improve deployment safety.
  • Maintain production standards, including runbooks, monitoring, alerting, change management, and deployment readiness reviews.
  • Construct, maintain, and optimize data centers and specialized AI infrastructure.
  • Advise application and infrastructure teams on reliability and mentor junior SREs.

Requirements

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • At least 5 years of experience in Site Reliability Engineering, Production Engineering, or a similar role.
  • Strong proficiency in a programming or scripting language such as Go, Python, or Bash.
  • Deep understanding of Linux/Unix operating systems, TCP/IP, DNS, networking fundamentals, and distributed systems.
  • Preferred experience managing large-scale data infrastructure such as MySQL, Redis, Kafka, and Flink.
  • Preferred production experience with Kubernetes and container orchestration.
  • Preferred expertise in designing, analyzing, and troubleshooting large-scale distributed systems.
  • Preferred experience leading incident response for complex, high-impact events and operating or constructing data centers.

Benefits

  • Rotational on-call schedule providing nonstop coverage for critical data infrastructure.
  • Collaboration across multiple time zones with opportunities to work alongside senior engineers.
  • Hands-on experience with modern infrastructure technologies, SRE practices, large-scale data systems, and AI infrastructure.

Tech Stack

Apache FlinkApache KafkaBashGoKubernetesLinuxMySQLPythonRedis

Categories

Site Reliability
ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me