8 days ago
Washington, DC, USAMid Level
Responsibilities
- Support and troubleshoot REST APIs, including HTTP requests and responses, authentication, authorization, JSON payloads, connectivity, rate limits, and API gateway errors.
- Provide Tier 2 technical support for API consumer and producer issues, including outages, access problems, deployment issues, and unexpected platform behavior.
- Perform platform operations, routine maintenance, deployments, configuration changes, troubleshooting, and recovery across AWS and Linux environments.
- Develop and maintain CI/CD pipelines using Jenkins, Harness, GitHub Actions, or comparable tools.
- Use Terraform and Ansible for infrastructure-as-code, configuration management, repeatable platform configuration, and controlled infrastructure changes.
- Use Docker, Kubernetes, and Helm for containerized deployments, release management, scaling, service exposure, ingress, and load balancing.
- Troubleshoot networking, load balancers, security groups, connectivity, high availability, and Apigee platform configurations.
- Develop Python, Node.js, or shell scripts and lightweight automations to reduce repetitive support and infrastructure work.
- Identify root causes of recurring operational problems and recommend documentation, process, automation, or platform improvements.
- Develop runbooks, troubleshooting guidance, standard operating procedures, and technical documentation.
- Participate in API onboarding and validation demonstrations and communicate technical findings to technical and non-technical stakeholders.
Requirements
- U.S. citizenship is required to obtain Public Trust; DHS USCIS Public Trust is preferred.
- At least three years of related platform engineering experience is required.
- A bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field and at least four years of relevant experience, or an equivalent combination of education and experience consistent with the applicable contract LCAT.
- Hands-on experience supporting and troubleshooting REST APIs and experience with an API gateway or enterprise API management platform; Apigee is strongly preferred.
- Working knowledge of AWS and experience operating or troubleshooting production or production-like cloud workloads.
- Working knowledge of Linux, preferably RHEL, including package management, permissions, security fundamentals, logs, shell scripting, and command-line troubleshooting.
- Experience with operational maintenance, deployment support, incident investigation, CI/CD tools, Python, Node.js or shell scripting, Terraform, and Ansible.
- Hands-on experience with Helm and familiarity with Docker and Kubernetes-based platform operations.
- Independent experience with either Terraform-based infrastructure-as-code or Docker, Kubernetes, and Helm-based deployments.
- Knowledge of networking, load balancers, high availability, DNS, security groups, and network troubleshooting.
- Ability to investigate technical problems, document findings, recommend corrective actions, and collaborate across technical, support, research, leadership, and government stakeholder teams.
- Ability and willingness to learn Power Apps, Power Automate, or comparable lightweight automation tools.
- Preferred qualifications include AWS services experience, technical documentation and runbook development, federal program experience, and certifications such as RHCSA, AWS, Kubernetes, or Terraform.
