Shield AI

Sr. Staff Site Reliability Engineer (R5803)

Shield AI
Apply
22 hours ago
San Mateo, CA, USAStaff+

Base Salary

$220k - $330k/yr

Responsibilities

  • Define and implement SLIs, SLOs, and other service reliability measures.
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • Lead technical response to complex incidents and drive root-cause analysis through resolution.
  • Identify recurring failure modes and partner with engineering teams to eliminate them.
  • Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • Develop tooling and automation to reduce manual operational work.
  • Partner with product and platform teams to incorporate reliability requirements into system design.
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery.
  • Mentor product engineers and Cloud Engineering and Reliability teammates on operational practices.
  • Define and manage short- and long-term SRE roadmaps and distribute work across teammates.

Requirements

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • Experience designing and operating infrastructure in AWS or another major cloud environment.
  • Experience with infrastructure-as-code and automated infrastructure provisioning.
  • Experience supporting containerized applications and distributed systems.
  • Experience developing operational tooling or automation using Python, Go, or a similar language.
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
  • Experience leading incident response and root-cause analysis across engineering teams.
  • Experience leading and executing the technical vision of a team over multi-quarter timelines.
  • Preferred: experience establishing or maturing an SRE function, using Kubernetes and cloud-native observability systems, and operating in regulated or compliance-driven environments.
  • Preferred: experience with capacity planning, performance analysis, cloud cost management, shared infrastructure, ticketing-system roadmaps, and engineering mentorship.

Benefits

  • Full-time regular employee package includes bonus, benefits, and equity.
  • Temporary employees receive a temporary benefits package applicable after 60 days of employment.
  • Offers are contingent on a cleared background and possible reference check.
  • Military fellows and part-time employees are not eligible for benefits.

Categories

Site Reliability
Shield AI

About Shield AI

1,001-5,000 employees

Founded in 2015, Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide.

Contact me