20 days ago
Base Salary
$104k - $166k/yr
Responsibilities
- Operate and maintain production infrastructure services and applications for availability, reliability, performance, security, and operational health.
- Monitor services using SLIs, SLOs, dashboards, alerts, metrics, logs, traces, and observability tooling.
- Manage incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and corrective actions.
- Execute application and infrastructure releases through staging and production deployment pipelines, including validation and rollback.
- Manage infrastructure upgrades, patching, configuration changes, maintenance, and technology refreshes.
- Improve resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
- Identify reliability risks and technical debt using reliability metrics, incident trends, capacity data, and service health indicators.
- Automate operational activities using infrastructure-as-code practices and collaborate with platform and application engineering teams.
Requirements
- U.S. citizenship and ability to obtain and maintain the required Public Trust level clearance.
- Bachelor’s degree and eight years of experience, or a high school diploma/equivalent and twelve years of experience.
- At least seven years of hands-on experience in site reliability engineering, DevOps, or production systems engineering.
- Hands-on experience operating AWS Commercial and AWS GovCloud, including OpenShift ROSA or comparable Kubernetes-based platforms.
- Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
- Experience with GitLab and Jenkins CI/CD platforms, including reliability gating and deployment automation.
- Proficiency administering Linux and Windows Server.
- Experience with Dynatrace, Datadog, Splunk, and OpenTelemetry.
- Experience owning SLI/SLO and alerting programs, including error budgets, alert rationalization, and noise reduction.
- Scripting or automation proficiency in Python, Bash, PowerShell, or Go.
- Experience operating in federal or regulated environments such as FISMA, FedRAMP, and NIST 800-53.
- Preferred qualifications include AWS, Red Hat, Azure, GCP, Dynatrace, Datadog, GitLab, Jenkins, and Terraform certifications.
Benefits
- Remote work arrangement.
- Evening shift schedule from 3pm to 11pm Eastern Standard Time.
- Employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.
Tech Stack
AnsibleAWSAzureBashDatadogGoGoogle Cloud PlatformJenkinsKubernetesLinuxOpenShiftPowerShellPythonSplunkTerraform
Categories
Site Reliability
About Peraton
Peraton builds and integrates mission systems, cyber, space, and intelligence solutions for U.S. defense, intelligence, and civil agencies. It delivers enterprise IT, satellite and terrestrial communications, and spectrum management under government contracts. Founded in 2017, the company is headquartered in Reston, Virginia and is privately held by Veritas Capital.
