2 days ago
San Jose, CA, USAStaff+
Base Salary
$124k - $271k/yr
Responsibilities
- Own and improve SLO/SLI practices for real-time services, including latency, availability, jitter, and packet loss metrics.
- Lead critical incident response, postmortems, chaos engineering, and game day exercises across the real-time platform.
- Build and evolve dashboards, alerting, distributed tracing, runbooks, and architectural decision records for media infrastructure.
- Act as an architectural authority for deployment patterns, infrastructure design, scalability, fault tolerance, and operational readiness.
- Drive capacity planning, traffic modeling, cost optimization, infrastructure tooling evaluation, and deployment safety across globally distributed systems.
- Partner with multiple engineering teams on planning, reliability guidance, CI/CD, infrastructure-as-code, GitOps, testing, and deployment automation.
- Guide senior engineers on SRE principles and reliability patterns while coordinating with engineering, product, operations, networking, security, and data teams.
- Serve as a technical liaison across US, China, and India teams, including architecture reviews and planning in English and Mandarin as appropriate.
Requirements
- 10+ years of experience in DevOps, SRE, or infrastructure engineering, including at least 3 years at staff or principal scope.
- Proven ownership of reliability for large-scale, distributed, latency-sensitive production systems.
- Experience supporting real-time or media-heavy platforms such as video conferencing, live streaming, gaming, or trading systems.
- Ability to lead cross-functional technical initiatives without direct authority and drive alignment across engineering, product, and operations.
- Architectural understanding of WebRTC, RTP/RTCP, TURN/STUN, SDP, and SFU/MCU topologies.
- Expertise with cloud infrastructure and container orchestration, including AWS, GCP or Azure, Kubernetes, Helm, and ArgoCD.
- Experience with infrastructure-as-code tools such as Terraform or Pulumi.
- Experience with observability tools such as Prometheus, Grafana, Datadog, Jaeger, or OpenTelemetry.
- Understanding of BGP, anycast routing, DNS, load balancing, and CDN architecture.
- Experience with GitHub Actions, Jenkins, and Spinnaker, as well as canary, feature-flag, and blue/green deployment strategies.
- Proficiency in Python, Bash, or Go for automation, tooling, and incident response.
- Ability to work across global time zones; occasional weekend work may be required.
Benefits
- Structured hybrid work model combining office and remote environments.
- Benefits and perks supporting physical, mental, emotional, and financial health, work-life balance, and community involvement.
- Flexible schedule and global team collaboration across Beijing, Shanghai, Bangalore, Hyderabad, and US time zones.
- Potential closing date is October 2, 2026.
Tech Stack
AWSAzureBashDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesPrometheusPythonSpinnakerTerraform
Categories
DevOpsSite Reliability
About Zoom
Zoom builds a cloud platform for video meetings, team chat, webinars, phone, rooms, and contact center used by businesses, schools, and public-sector organizations. It sells these collaboration and communications services as SaaS subscriptions with add-ons for events, telephony, and enterprise features, and offers developer APIs/SDKs. Founded in 2013 and headquartered in San Jose, California, Zoom Video Communications is a public company traded on Nasdaq.
