4 months ago
Arlington, VA, USASenior
Base Salary
$192k - $455k/yr
Responsibilities
- Define, track, and report on SLIs, SLOs, and error budgets for platform services in the customer environment.
- Identify and automate operational toil and translate reliability data into prioritized engineering work.
- Own on-site observability, including dashboards, alerting thresholds, and log pipelines using the LGTM stack.
- Lead incident response, triage, containment, coordination, customer communication, post-incident reviews, root cause analysis, and preventive fixes.
- Maintain runbooks, emergency response procedures, and on-call escalation processes.
- Deploy, update, configure, and roll back containerized services using Docker and Docker Compose.
- Apply and validate Terraform infrastructure changes within DSO policy guardrails.
- Perform capacity planning and communicate scaling requirements to the Arlington engineering team.
- Act as the primary technical interface between the government customer and Twenty’s engineering team.
- Provide technical troubleshooting guidance and partner on compliance, logging, and audit requirements.
Requirements
- 5+ years of professional experience in site reliability engineering, production operations, or a closely related infrastructure role.
- Production experience defining and tracking SLIs, SLOs, and error budgets.
- Hands-on production experience with Docker, Docker Compose, and AWS services including EC2, ECS, RDS, VPCs, and security groups.
- Solid Linux/Unix systems administration skills and experience working in constrained environments.
- Experience with Terraform for infrastructure provisioning and configuration within policy guardrails.
- Experience with the LGTM observability stack or equivalent, including Grafana, Loki, Prometheus/Mimir, and distributed tracing.
- Strong incident response experience, including leading responses, writing post-mortems and runbooks, and implementing preventive fixes.
- Scripting proficiency in Python or Bash; familiarity with Go and PagerDuty or equivalent on-call tooling.
- Experience supporting government or defense environments, including air-gapped or enclave deployments.
- Active TS/SCI security clearance with appropriate polygraph and the ability to maintain it.
- U.S. citizenship and willingness to travel occasionally for customer engagements and operational support.
- Nice-to-have experience with NATS or similar pub/sub systems, cyber operations or intelligence environments, and AWS certifications such as Solutions Architect, SysOps, or DevOps Engineer.
Benefits
- Medical, dental, and vision plan options, plus life/AD&D and disability coverage options.
- Paid parental leave for eligible full-time employees, including 12 weeks for birthing parents, 4 weeks for non-birthing parents, and 6 weeks for adoptive, foster, or intended parents through surrogacy.
- Paid holidays and flexible PTO.
- 401(k) with pre-tax and Roth options, HSA/FSA options, and dependent care FSA.
- On-site, full-time work at a government customer site with occasional travel for customer engagements and operational support.
Categories
Forward DeployedSite Reliability