ByteDance

Site Reliability Engineer (Multiple Positions)

ByteDance
Apply
20 hours ago
San Jose, CA, USAMid Level
H1B sponsor

Base Salary

$213k - $388k/yr

Responsibilities

  • Design and develop highly available infrastructure platforms supporting large-scale distributed production systems.
  • Build automation systems for infrastructure provisioning, deployment orchestration, configuration management, and operational workflows.
  • Monitor system performance, availability, and reliability; identify bottlenecks, capacity risks, and reliability issues.
  • Develop reliability engineering frameworks for system stability, fault tolerance, and high-load and failure scenarios.
  • Design and implement disaster recovery architectures, failover strategies, and business continuity solutions across multi-region environments.
  • Oversee incident response, conduct root cause analysis, and implement long-term remediation strategies.
  • Participate in technical operations and on-call rotations for performance and reliability issues.
  • Mentor junior Site Reliability Engineers.

Requirements

  • Master’s degree or foreign equivalent in Computer Science, Engineering, Information Technology, Mathematics, or a related field plus two years of related work experience; or a bachelor’s degree or foreign equivalent in one of those fields plus five years of post-bachelor’s progressive related experience.
  • At least two years of experience across the software development lifecycle for back-end and cloud-native projects.
  • At least two years of experience designing scalable and highly available software systems for cloud environments.
  • At least two years of experience processing and analyzing large quantities of logs and data, and creating reports, dashboards, and alerts using search processing languages.
  • At least two years of Linux administration experience, including performance monitoring, debugging, and troubleshooting networked applications.
  • At least two years of experience developing Docker containers and deploying, managing, and monitoring services in Kubernetes clusters.

Benefits

  • Full-time position requiring 40 hours per week.
  • On-site location in San Jose, California.

Categories

Site Reliability
ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me