22 hours ago
San Diego, CA, USAStaff+
Base Salary
$183k - $275k/yr
Responsibilities
- Define and implement SLIs, SLOs, and other service reliability measures.
- Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
- Lead technical response to complex incidents and drive root-cause analysis through resolution.
- Identify recurring failure modes and partner with engineering teams to eliminate them.
- Improve system resilience through automation, testing, capacity planning, and failure recovery.
- Develop tooling and automation to reduce manual operational work.
- Incorporate reliability requirements into product and platform system designs.
- Establish incident response practices that improve detection, diagnosis, communication, and recovery.
- Mentor product engineers and SRE and Cloud Engineering teammates on reliability and operational practices.
- Define and manage short- and long-term SRE roadmaps and distribute work across teammates.
Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
- Experience operating production services with defined availability and reliability requirements.
- Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
- Experience designing and operating infrastructure in AWS or another major cloud environment.
- Experience with infrastructure-as-code and automated infrastructure provisioning.
- Experience supporting containerized applications and distributed systems.
- Experience developing operational tooling or automation using Python, Go, or a similar language.
- Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
- Experience leading incident response and root-cause analysis across engineering teams.
- Experience leading and executing a team’s technical vision over multi-quarter timelines.
- Preferred experience establishing or maturing an SRE function, using Kubernetes and cloud-native observability systems, and operating in regulated or compliance-driven environments.
- Preferred experience with capacity planning, performance analysis, cloud cost management, shared infrastructure, ticketing-system roadmaps, and engineering mentorship.
Benefits
- Full-time regular employees receive a bonus, benefits, and equity in addition to pay within the listed range.
- Temporary employees receive a temporary benefits package applicable after 60 days of employment.
- Offers are contingent on a cleared background and possible reference check.
- Military fellows and part-time employees are not eligible for benefits.
Tech Stack
Categories
DevOpsSite Reliability
About Shield AI
Founded in 2015, Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide.
