1 month ago
Washington, DC, USASenior
Base Salary
$113k - $162k/yr
Responsibilities
- Lead, mentor, and inspire a team of Site Reliability Engineers supporting Managed Gateway offerings.
- Architect and implement scalable, fault-tolerant cloud-native systems across AWS, GCP, and Azure.
- Own monitoring, alerting, incident response, post-mortems, SLOs, SLIs, and continuous reliability improvements.
- Build automation, self-service tooling, and workflows for deploying and managing API gateways.
- Drive technical debt prevention, architectural best practices, and operational readiness for new features.
- Partner with Product, Engineering, Customer Success, and Professional Services on roadmap and implementation decisions.
- Lead enterprise Cloud Gateway onboarding from kickoff through successful production implementation.
- Handle complex customer topologies, serve as the technical implementation owner, and escalate technically complex accounts.
- Productize recurring implementation patterns into playbooks and platform capabilities.
- Feed customer implementation constraints and patterns back into Product.
Requirements
- Extensive experience as a Site Reliability Engineer working with highly available and distributed systems.
- Deep expertise with Kubernetes and cloud-native architectures across multiple public cloud providers, preferably AWS, GCP, and Azure.
- Strong proficiency in Golang or similar modern programming languages for automation and tool development.
- Experience building and maintaining CI/CD pipelines and infrastructure as code with Terraform and Ansible.
- Knowledge of monitoring, logging, and alerting systems such as Prometheus, Grafana, ELK stack, and Datadog.
- Experience with managed services, API gateways, or similar network infrastructure is highly desirable.
- Experience with service mesh technologies such as Istio or Linkerd is a bonus.
- Familiarity with PostgreSQL or Cassandra administration for high-throughput systems is a bonus.
- Contributions to open-source SRE tools or projects are a bonus.
- Relevant cloud certifications such as AWS Certified DevOps Engineer or CKA are a bonus.
- Strong ownership, urgency during critical incidents, collaboration, and adaptability are expected.
Tech Stack
AnsibleApache CassandraAWSAzureDatadogGoGoogle Cloud PlatformGrafanaIstioKubernetesPostgreSQLPrometheusTerraform
Categories
DevOpsSite Reliability
