Base Salary
$104k - $166k/yr
Responsibilities
- Operate and maintain production infrastructure services and applications for availability, reliability, performance, security, and operational health.
- Monitor systems with SLIs, SLOs, dashboards, alerts, metrics, logs, traces, and observability tooling.
- Manage incidents and service disruptions through on-call response, troubleshooting, restoration, root-cause analysis, and corrective actions.
- Execute application and infrastructure releases through staging and production, including validation, rollback, and release troubleshooting.
- Manage upgrades, patching, configuration changes, maintenance, and technology refreshes across deployed infrastructure.
- Improve resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
- Automate operational activities using infrastructure-as-code practices and improve reusable infrastructure building blocks with platform and application teams.
Requirements
- U.S. citizenship and the ability to obtain and maintain the required Public Trust clearance.
- Bachelor’s degree and eight years of experience, or a high school diploma/equivalent and twelve years of experience.
- At least seven years of hands-on experience in site reliability engineering, DevOps, or production systems engineering.
- Hands-on experience operating AWS Commercial and AWS GovCloud environments, including ROSA/OpenShift or comparable Kubernetes platforms.
- Strong infrastructure-as-code experience with Terraform and Ansible or Ansible Tower.
- Experience with GitLab and Jenkins CI/CD platforms, including reliability gating and deployment automation.
- Proficiency administering Linux and Windows Server.
- Experience with Dynatrace, Datadog, Splunk, and OpenTelemetry or comparable enterprise observability tools.
- Ownership of SLI/SLO and alerting programs, including error budgets, alert rationalization, and noise reduction.
- Scripting or automation proficiency in Python, Bash, PowerShell, or Go.
- Experience operating in federal or regulated environments such as FISMA, FedRAMP, or NIST 800-53.
- Preferred qualifications include AWS, Red Hat, Azure, GCP, Dynatrace, Datadog, GitLab, Jenkins, or Terraform certifications.
Benefits
- Remote work arrangement.
- Day-shift schedule from 7 a.m. to 3 p.m. Eastern Standard Time.
- Employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.
Tech Stack
Categories
About Peraton
At Peraton, we're at the forefront of delivering the next big thing every day. We're the partner of choice to help solve some of the world's most daunting challenges, delivering bold, new solutions to keep people around the world safer and more secure. How do we do it? By thinking differently. We're not mired in the past. We look at all problems with fresh eyes. We look past the obvious to bring the best talent, tech, and ideas together to completely transform how things get done. So bring your unique ideas, your entrepreneurial spirit, and your drive to succeed and get ready to be part of something bigger. Get ready to do the can't be done. ________ Recruitment fraud is a growing trend where fraudsters have been known to attempt to use our name to trick job seekers with fake employment opportunities. This type of scam is typically carried out through fake job postings, fake websites, or email accounts claiming to be from Peraton. The intent of recruitment fraud is to gain access to your personal information, such as your banking information, credit card number, or social security number. Please be aware that our careers site can be found at careers.peraton.com and our corporate site can be found at peraton.com. To learn more about Recruitment fraud and what to expect and not to expect from a Peraton recruiter, please visit: https://careers.peraton.com/recruitment-fraud/