Box

Senior Technical Duty Officer, Cloud Ops

Box
Apply
13 days ago
Redwood City, CA, USASenior
H1B sponsor

Base Salary

$187k - $234k/yr

Responsibilities

  • Own and direct critical, blocker, and other high-severity production incidents through mitigation and recovery.
  • Lead incident bridges by clarifying impact, coordinating SMEs and resources, delegating work, and driving rapid service restoration.
  • Improve incident-platform tooling by automating repetitive response steps and developing reusable operational tools.
  • Partner with SRE and engineering teams to improve service manageability, observability, resiliency, and operational readiness.
  • Lead daily change reviews and assess planned-change risk with engineering teams.
  • Represent the GTOC/NOC in problem management, readiness reviews, and reliability forums.
  • Use observability APIs, PagerDuty, Jira, and internal services to make response workflows measurable and repeatable.
  • Mentor team members through training, runbooks, tabletop exercises, and operational process improvements.

Requirements

  • 5+ years of experience in SRE, production operations, reliability engineering, or equivalent high-scale internet or SaaS operations, including repeated major-incident leadership.
  • Demonstrated Incident Commander, Technical Duty Officer, or equivalent experience with escalation judgment, delegation, and communication under uncertainty.
  • Strong SRE knowledge covering SLIs, SLOs, error budgets, observability, golden signals, blameless postmortems, toil reduction, and product-team operational ownership.
  • Proficiency in Python for maintainable automation and tooling, including APIs, packaging or service-style tools, and code review.
  • Solid Linux/Unix troubleshooting skills and experience with multi-tier distributed systems.
  • Networking knowledge covering DNS, TLS, load balancing, HTTP, and basic routing and firewall concepts.
  • Experience with cloud environments, preferably GCP, as well as Kubernetes or equivalent container-orchestration systems.
  • Excellent written and verbal communication, including executive-ready impact statements, incident-bridge facilitation, and trusted documentation.
  • Ability to coach and uplift others through mentoring, training, drills, and runbook improvement.
  • Preferred qualifications include 24x7 NOC/GTOC experience, Prometheus-compatible observability, distributed tracing, synthetic monitoring, PagerDuty, incident tooling, emergency change management, Go, shell, Terraform, CI/CD, dependency graphs, service catalogs, or Tier 1 journey documentation.

Benefits

  • Eligible for equity and benefits.
  • Assigned-office work is required at least three days per week.
  • Box provides reasonable accommodations for applicants with disabilities and emphasizes an inclusive, diverse workplace.

Categories

Site Reliability
Box

About Box

1,001-5,000 employees

Box (NYSE:BOX) is the Intelligent Content Cloud, a single platform that enables organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI. Founded in 2005, Box simplifies work for leading global organizations, including JLL, Morgan Stanley, and Nationwide. Box is headquartered in Redwood City, CA, with offices across the United States, Europe, and Asia. Visit box.com to learn more. And visit box.org to learn more about how Box empowers nonprofits to fulfill their missions.

Contact me