3 days ago
Base Salary
$99k - $229k/yr
Responsibilities
- Own the SLO/SLI framework and improve reliability metrics including latency, availability, jitter, and packet loss for real-time services.
- Lead critical incident response, postmortems, chaos engineering, game days, and measurable reliability improvements.
- Build observability capabilities including dashboards, alerting, and distributed tracing for real-time media infrastructure.
- Serve as an architectural authority for deployment patterns, infrastructure design, operational readiness, scalability, and fault tolerance.
- Drive capacity planning, traffic modeling, infrastructure cost optimization, and evaluation of infrastructure tools, platforms, vendors, media servers, CDN providers, and edge networking.
- Establish standards for CI/CD, infrastructure-as-code, GitOps, deployment automation, canary releases, feature flags, and blue/green deployments.
- Partner with engineering teams on operational readiness and guide senior engineers on SRE principles and reliability patterns.
- Collaborate across network engineering, security, product, and data teams and provide technical liaison support across US, China, and India teams.
- Conduct architecture reviews, incident retrospectives, and planning sessions in English and Mandarin as appropriate.
- Maintain documentation, runbooks, and architectural decision records for global team members.
- Participate in an on-call rotation and maintain a schedule that overlaps with teams in Beijing, Shanghai, Bangalore, and Hyderabad.
Requirements
- 5+ years of experience in DevOps, SRE, or infrastructure engineering, including at least 3 years at staff or principal level scope.
- Proven ownership of reliability for large-scale, distributed, latency-sensitive production systems.
- Experience supporting real-time or media-heavy platforms such as video conferencing, live streaming, gaming, or trading systems.
- Ability to lead cross-functional technical initiatives without direct authority and align engineering, product, and operations teams.
- Conceptual and architectural understanding of WebRTC, RTP/RTCP, TURN/STUN, SDP, and SFU/MCU topologies.
- Solid expertise with cloud infrastructure using AWS, GCP, or Azure.
- Experience with infrastructure-as-code tooling such as Terraform or Pulumi.
- Experience with observability tools including Prometheus, Grafana, Datadog, Jaeger, or OpenTelemetry.
- Understanding of BGP, anycast routing, DNS, load balancing, and CDN architecture.
- Experience with CI/CD tools such as GitHub Actions, Jenkins, and Spinnaker.
- Proficiency in Ansible, Python, Bash, or Go for automation, tooling, and incident response.
- Ability to work globally or across multiple time zones.
- Mandarin-speaking ability is preferred.
- US citizenship is preferred.
Benefits
- Structured hybrid work approach centered around office and remote-work environments.
- Benefits and perks supporting physical, mental, emotional, and financial health, work-life balance, and community involvement.
- Application window is at least five days, with an anticipated position close date of October 1, 2026.
Tech Stack
Categories
DevOpsSite Reliability
About Zoom
Zoom builds a cloud platform for video meetings, team chat, webinars, phone, rooms, and contact center used by businesses, schools, and public-sector organizations. It sells these collaboration and communications services as SaaS subscriptions with add-ons for events, telephony, and enterprise features, and offers developer APIs/SDKs. Founded in 2013 and headquartered in San Jose, California, Zoom Video Communications is a public company traded on Nasdaq.
