
Technical Site Reliability Engineer
Anduril IndustriesResponsibilities
- Maintain the simulation software stack, including installation, configuration, updates, version management, and daily functionality.
- Own compute, networking, storage, and environment configuration infrastructure; keep it provisioned, patched, and performant.
- Build and maintain automated post-release regression and smoke-test suites for software releases and configuration changes.
- Forecast, diagnose, and eliminate failure modes through root-cause analysis, monitoring, guardrails, and process improvements.
- Partner with development teams and stakeholders to identify reliability risks and implement mitigation strategies before releases.
- Monitor system health, instrument the environment, triage issues, and escalate problems with actionable context.
- Document runbooks, known issues, environment configuration, and release-validation results.
Requirements
- Proficiency in Python for automation, tooling, and test development.
- Working knowledge of C++ sufficient to read, debug, build, and trace issues in simulation software.
- Knowledge of TCP/IP, UDP, multicast, DNS, routing, firewalls, latency, packet loss, and distributed-system connectivity diagnostics.
- Experience maintaining production or production-adjacent systems and troubleshooting under time pressure.
- Experience with project management, issue tracking, bug triage, and coordination across engineering teams.
- Strong written and verbal communication skills for escalation, root-cause explanation, and runbook creation.
- Eligibility to pass security and background-check requirements for sensitive information systems.
- Preferred experience with modeling and simulation, wargaming, or distributed simulation standards such as DIS, HLA, and TENA.
- Preferred experience with platforms such as AFSIM, VBS, or similar tools.
- Preferred test automation and CI/CD experience, including building automated validation pipelines.
- Preferred infrastructure-as-code and configuration-management experience with Terraform, Ansible, Docker, or Kubernetes.
- Preferred on-premises and cloud deployment experience.
- Preferred experience with observability tooling such as Prometheus, Grafana, or ELK.
- Preferred depth in Linux systems administration and comfort working in mixed Linux/Windows environments.
- Prior defense, aerospace, or classified-environment experience and an active security clearance are preferred.
Benefits
- Comprehensive benefits package available at little to no cost to full-time employees.
- The position is in Abu Dhabi, UAE, with initial hiring in London, UK, and requires relocation to the facility after completion.
- Highly competitive equity grants are included in the majority of full-time offers.
- Full-time employee benefits support health and recovery.
Categories
About Anduril Industries
Anduril is not a traditional defense contractor. We are shaping the future of defense, transforming US & allied military capabilities with advanced technology. We emphasize speed and results and control our products from start to finish, including funding R&D to selling finished products off the shelf. Today, Anduril is in a rapid growth phase, deploying technology in diverse locations and developing path-making products that will change defense forever. We believe that everyone at Anduril can be a catalyst. Your perspective can change lives, and we want to help you make your mark. Our team includes thinkers and doers working interdependently. We bring the brightest minds and best-in-class talent together with veterans who have lived the problems of our warfighters. If you like building quickly and seeing your work deployed in the real world, we want you at Anduril. With offices in Orange County, Washington DC, Seattle, Boston, Atlanta, London, and Sydney, our reach is wide. Check out our careers page at https://www.anduril.com/careers or send us an email at careers@anduril.com.