5 months ago
Base Salary
$180k - $440k/yr
Responsibilities
- Develop and optimize software to provision and manage infrastructure across on-premise, virtual machine, and classified cloud environments.
- Improve infrastructure reliability, performance, and cost-effectiveness for large-scale AI and application workloads.
- Collaborate with engineers to design solutions for government-specific workloads and compliance requirements.
- Implement observability, monitoring, and security practices to protect the integrity, availability, and confidentiality of critical systems.
- Manage secure storage infrastructure using infrastructure-as-code tools such as Pulumi, Terraform, or Ansible.
- Lead incident management and postmortems and define clear SLAs and SLOs to drive system reliability.
- Support infrastructure across bare-metal, classified cloud, and hybrid cloud architectures.
Requirements
- Active Top Secret security clearance is required.
- At least five years of experience as an Infrastructure Engineer, Site Reliability Engineer, or in a similar role building and maintaining reliable, scalable systems.
- Proficiency managing storage infrastructure with Pulumi, Terraform, or Ansible.
- Deep understanding of the Kubernetes stack, including CNI, CRI, CSI, and related components.
- Experience improving system reliability through incident management, postmortems, and defining SLAs and SLOs.
- Strong communication and documentation skills, including the ability to handle sensitive information concisely and accurately.
- Preferred experience installing, configuring, debugging, and maintaining GPU hardware and drivers.
- Preferred experience supporting high-traffic web or mobile workloads and optimizing Kubernetes for large-scale classified or federal deployments.
- Familiarity with chaos engineering, capacity planning, or similar resilience practices is preferred.
- Proficiency with Kyverno, ArgoCD, or Go for infrastructure automation is preferred.
- Security certifications such as CISSP or experience in secure federal environments are preferred.
Benefits
- In-person role based in Palo Alto, California, or Washington, DC.
- Up to 50% travel required.
- Equity, medical, vision, and dental coverage.
- 401(k) retirement plan.
- Short- and long-term disability insurance and life insurance.
- Various discounts and perks.
Tech Stack
Categories
DevOpsSite Reliability
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers