3 hours ago
Responsibilities
- Own campus-scale monitoring architecture, alert suppression, and signal quality.
- Provide technical incident command, bridge coordination, and severity and timeline management for SEV events.
- Run blameless postmortems and drive corrective actions through completion.
- Lead cross-functional reliability projects across compute, network, storage, power, cooling, and facility signal boundaries.
- Build and maintain playbooks, run game days, and keep dependency maps current with the NOC.
- Define error budgets and availability objectives at campus and service boundaries.
- Participate in on-call rotations and SEV incident response at the Memphis/Southaven data center campus.
Requirements
- Bachelor’s degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least five years of experience in site reliability, systems engineering, or large-scale production operations.
- Large-scale incident command experience and calm technical leadership during incidents.
- Fleet- or campus-scale monitoring and observability design experience, including alert hygiene, suppression, and signal quality.
- Experience spanning at least two of compute, network, storage, power, and cooling or facilities telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- Proficiency in Python and Bash scripting plus general experience with at least one systems language such as C, C++, Java, Go, or Rust.
- Strong problem-solving skills with a data-driven reliability engineering approach and the ability to collaborate across NOC, data center operations, and infrastructure engineering.
- Preferred experience includes AI/ML infrastructure or supercomputing, SLO/SLI and error budget definition, game days, dependency mapping, closed-loop corrective actions, data center hardware and plant signals, and startup or technology-company experience.
Benefits
- On-call rotation and incident response responsibilities are part of the role.
- Work is based at the Memphis/Southaven data center campus.
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers