Shield AI

Sr. Staff Site Reliability Engineer (R5803)

Shield AI
Apply
22 hours ago
San Diego, CA, USAStaff+

Base Salary

$183k - $275k/yr

Responsibilities

  • Define and implement SLIs, SLOs, and other service reliability measures.
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • Lead technical response to complex incidents and drive root-cause analysis through resolution.
  • Identify recurring failure modes and partner with engineering teams to eliminate them.
  • Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • Develop tooling and automation to reduce manual operational work.
  • Incorporate reliability requirements into product and platform system designs.
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery.
  • Mentor product engineers and SRE and Cloud Engineering teammates on reliability and operational practices.
  • Define and manage short- and long-term SRE roadmaps and distribute work across teammates.

Requirements

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • Experience designing and operating infrastructure in AWS or another major cloud environment.
  • Experience with infrastructure-as-code and automated infrastructure provisioning.
  • Experience supporting containerized applications and distributed systems.
  • Experience developing operational tooling or automation using Python, Go, or a similar language.
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
  • Experience leading incident response and root-cause analysis across engineering teams.
  • Experience leading and executing a team’s technical vision over multi-quarter timelines.
  • Preferred experience establishing or maturing an SRE function, using Kubernetes and cloud-native observability systems, and operating in regulated or compliance-driven environments.
  • Preferred experience with capacity planning, performance analysis, cloud cost management, shared infrastructure, ticketing-system roadmaps, and engineering mentorship.

Benefits

  • Full-time regular employees receive a bonus, benefits, and equity in addition to pay within the listed range.
  • Temporary employees receive a temporary benefits package applicable after 60 days of employment.
  • Offers are contingent on a cleared background and possible reference check.
  • Military fellows and part-time employees are not eligible for benefits.

Categories

DevOpsSite Reliability
Shield AI

About Shield AI

1,001-5,000 employees

Founded in 2015, Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide.

Contact me