Technical Site Reliability Engineer
Anduril IndustriesResponsibilities
- Maintain the simulation software stack, including installation, configuration, updates, version management, and day-to-day functionality.
- Own and operate the compute, networking, storage, and environment configuration underlying the simulation systems.
- Design, automate, and continually extend post-release regression and smoke-test processes.
- Forecast, diagnose, and permanently resolve system failures and bugs while implementing monitoring and preventative guardrails.
- Review upcoming development changes for reliability risks and implement mitigation strategies before releases reach the simulation environment.
- Monitor system health, triage issues, and escalate problems with sufficient context for resolution.
- Document runbooks, known issues, environment configuration, and release-validation results.
Requirements
- Proficiency in Python for automation, tooling, and test development.
- Working knowledge of C++ sufficient to read, debug, build, and trace issues in the simulation codebase.
- Solid networking fundamentals, including TCP/IP, UDP, multicast, DNS, routing, and firewalls, with the ability to diagnose distributed-systems latency, packet loss, and connectivity problems.
- Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
- Demonstrated experience maintaining production or production-adjacent systems and troubleshooting under time pressure.
- Strong written and verbal communication skills, including explaining root causes, escalating issues, and writing runbooks.
- Eligibility to pass security and background-check requirements for sensitive information systems.
- Preferred experience with modeling and simulation, wargaming, or distributed simulation standards such as DIS, HLA, and TENA.
- Preferred experience with simulation platforms such as AFSIM, VBS, or similar.
- Preferred experience building automated validation pipelines and with infrastructure-as-code and configuration management.
- Preferred on-premises and cloud deployment experience.
- Preferred experience with Prometheus, Grafana, ELK, or equivalent observability tooling.
- Preferred depth in Linux systems administration and comfort working in mixed Linux/Windows environments.
- Prior defense, aerospace, or classified-environment experience and an active security clearance are preferred.
Benefits
- The position is based in Abu Dhabi, UAE, with initial hiring in London, UK, and requires relocation to the facility after completion.
- The role offers full-time benefits, including comprehensive health and recovery support available at little to no cost.
- Candidates must be eligible to pass security and background checks for sensitive information systems.
About Anduril Industries
Anduril is not a traditional defense contractor. We are shaping the future of defense, transforming US & allied military capabilities with advanced technology. We emphasize speed and results and control our products from start to finish, including funding R&D to selling finished products off the shelf. Today, Anduril is in a rapid growth phase, deploying technology in diverse locations and developing path-making products that will change defense forever. We believe that everyone at Anduril can be a catalyst. Your perspective can change lives, and we want to help you make your mark. Our team includes thinkers and doers working interdependently. We bring the brightest minds and best-in-class talent together with veterans who have lived the problems of our warfighters. If you like building quickly and seeing your work deployed in the real world, we want you at Anduril. With offices in Orange County, Washington DC, Seattle, Boston, Atlanta, London, and Sydney, our reach is wide. Check out our careers page at https://www.anduril.com/careers or send us an email at careers@anduril.com.