Zoom

Staff DevOps Engineer

Zoom
Apply
2 days ago
San Jose, CA, USAStaff+

Base Salary

$124k - $271k/yr

Responsibilities

  • Own and improve SLO/SLI practices for real-time services, including latency, availability, jitter, and packet loss metrics.
  • Lead critical incident response, postmortems, chaos engineering, and game day exercises across the real-time platform.
  • Build and evolve dashboards, alerting, distributed tracing, runbooks, and architectural decision records for media infrastructure.
  • Act as an architectural authority for deployment patterns, infrastructure design, scalability, fault tolerance, and operational readiness.
  • Drive capacity planning, traffic modeling, cost optimization, infrastructure tooling evaluation, and deployment safety across globally distributed systems.
  • Partner with multiple engineering teams on planning, reliability guidance, CI/CD, infrastructure-as-code, GitOps, testing, and deployment automation.
  • Guide senior engineers on SRE principles and reliability patterns while coordinating with engineering, product, operations, networking, security, and data teams.
  • Serve as a technical liaison across US, China, and India teams, including architecture reviews and planning in English and Mandarin as appropriate.

Requirements

  • 10+ years of experience in DevOps, SRE, or infrastructure engineering, including at least 3 years at staff or principal scope.
  • Proven ownership of reliability for large-scale, distributed, latency-sensitive production systems.
  • Experience supporting real-time or media-heavy platforms such as video conferencing, live streaming, gaming, or trading systems.
  • Ability to lead cross-functional technical initiatives without direct authority and drive alignment across engineering, product, and operations.
  • Architectural understanding of WebRTC, RTP/RTCP, TURN/STUN, SDP, and SFU/MCU topologies.
  • Expertise with cloud infrastructure and container orchestration, including AWS, GCP or Azure, Kubernetes, Helm, and ArgoCD.
  • Experience with infrastructure-as-code tools such as Terraform or Pulumi.
  • Experience with observability tools such as Prometheus, Grafana, Datadog, Jaeger, or OpenTelemetry.
  • Understanding of BGP, anycast routing, DNS, load balancing, and CDN architecture.
  • Experience with GitHub Actions, Jenkins, and Spinnaker, as well as canary, feature-flag, and blue/green deployment strategies.
  • Proficiency in Python, Bash, or Go for automation, tooling, and incident response.
  • Ability to work across global time zones; occasional weekend work may be required.

Benefits

  • Structured hybrid work model combining office and remote environments.
  • Benefits and perks supporting physical, mental, emotional, and financial health, work-life balance, and community involvement.
  • Flexible schedule and global team collaboration across Beijing, Shanghai, Bangalore, Hyderabad, and US time zones.
  • Potential closing date is October 2, 2026.

Tech Stack

AWSAzureBashDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesPrometheusPythonSpinnakerTerraform

Categories

DevOpsSite Reliability
Zoom

About Zoom

10,000+ employees

Zoom builds a cloud platform for video meetings, team chat, webinars, phone, rooms, and contact center used by businesses, schools, and public-sector organizations. It sells these collaboration and communications services as SaaS subscriptions with add-ons for events, telephony, and enterprise features, and offers developer APIs/SDKs. Founded in 2013 and headquartered in San Jose, California, Zoom Video Communications is a public company traded on Nasdaq.

Contact me