Box

Senior Incident Commander

Box
Apply
21 hours ago
Redwood City, CA, USASenior
H1B sponsor

Base Salary

$187k - $234k/yr

Responsibilities

  • Own and direct critical, blocker, and other high-severity live-site incidents through mitigation and recovery.
  • Triage incidents, clarify customer impact, organize incident bridges, delegate work, coordinate subject-matter experts, and restore service quickly.
  • Improve incident-platform tooling, templates, workflow automation, and response processes to reduce time to mitigate.
  • Partner with SRE and engineering teams to address operability gaps, dependencies, failure modes, service resiliency, manageability, and observability.
  • Lead daily reviews of planned changes in Jira and help evaluate and minimize change risk.
  • Provide technical leadership for incident, change, and problem-management capabilities in globally distributed 24x7 environments.
  • Translate incident learnings into engineering actions involving SLOs, signal quality, rollback readiness, and dependency documentation.
  • Use observability APIs, PagerDuty, Jira, and internal services to make response workflows measurable and repeatable.
  • Mentor team members through training, tabletop exercises, drills, and runbook improvements.
  • Represent GTOC/NOC in problem-management, readiness-review, and reliability forums.

Requirements

  • At least 5 years of experience in SRE, production operations, reliability engineering, or equivalent high-scale internet or SaaS operations.
  • Repeated experience leading or co-leading major production incidents as an Incident Commander, Technical Duty Officer, or equivalent.
  • Strong practical knowledge of SLIs, SLOs, error budgets, observability, golden signals, blameless postmortems, toil reduction, and product-team operations partnerships.
  • Proficiency in Python for maintainable automation and tooling, including APIs, packaging or service-style tools, and code review.
  • Experience troubleshooting Linux/Unix systems and working with multi-tier distributed systems spanning applications, data, networks, and cloud.
  • Networking knowledge covering DNS, TLS, load balancing, HTTP, and basic routing and firewall concepts.
  • Experience with cloud environments, preferably GCP, with AWS or Azure also valued.
  • Familiarity with Kubernetes or equivalent container and orchestration systems.
  • Strong written and verbal communication, including executive-ready impact statements, incident-bridge facilitation, and reliable documentation.
  • Ability to coach and uplift others through training, mentoring, runbook improvement, and drills.
  • Preferred experience in a 24x7 NOC, GTOC, or follow-the-sun operations center.
  • Preferred hands-on experience with Prometheus-compatible observability, Chronosphere, Grafana, SignalFx, Catchpoint, PagerDuty, ChatOps, workflow automation, status and communications templates, and drill frameworks.
  • Familiarity with emergency change management, break-glass procedures, CAB or asynchronous approval patterns, dependency graphs, service catalogs, and Tier 1 journey documentation.
  • Additional useful operations-automation experience includes Go, shell, Terraform, and CI/CD.

Benefits

  • Eligible for equity and benefits.
  • Expected to work from the assigned office at least 3 days per week.
  • Box provides reasonable accommodations for applicants with disabilities and promotes an inclusive workplace.

Categories

Site Reliability
Box

About Box

1,001-5,000 employees

Box builds a cloud content management platform for enterprises, combining secure file storage and sharing with collaboration, governance, e-signature, and workflow automation. It sells subscriptions (SaaS) to organizations that need to manage the lifecycle of business content and integrate with productivity apps. Founded in 2005 and headquartered in Redwood City, CA, Box is a public company (NYSE: BOX) used by customers such as JLL, Morgan Stanley, and Nationwide.

Contact me