10 days ago
Washington, DC, USAMid Level
Base Salary
$123k - $150k/yr
Responsibilities
- Operate and scale Kong’s multi-region, multi-tenant Konnect SaaS platform across AWS, GCP, and Azure.
- Build and maintain Kubernetes-based infrastructure and deployment workflows using Terraform, Terragrunt, Helm, and ArgoCD.
- Design and optimize multi-region PostgreSQL, Redis, ClickHouse, and Druid data and caching layers.
- Operate Kong Gateway and Kong Mesh environments supporting hybrid and distributed architectures.
- Develop CI/CD pipelines and GitOps workflows for automated service delivery and infrastructure changes.
- Improve observability and incident response using Datadog, Prometheus, Grafana, and Thanos while defining and tracking SLOs.
- Collaborate with development and security teams on reliable, secure, and compliant SaaS operations.
- Participate in a global 24/7 on-call rotation and improve operational playbooks and postmortem practices.
- Lead scaling initiatives that improve platform elasticity, reliability, resilience, and cost efficiency.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- Experience managing enterprise-scale SaaS or PaaS systems in multi-region, multi-tenant, secure environments.
- Deep Kubernetes expertise, including cluster and networking troubleshooting and fault-tolerant, scalable design.
- Strong proficiency with Terraform or Terragrunt and experience with CI/CD and GitOps workflows using ArgoCD, Atlantis, and Helm.
- Proficiency in one or more of Go, Python, or Bash for automation and tooling.
- Understanding of Linux/Unix, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.
- Experience with API gateway and service mesh technologies.
- Familiarity with Kafka and observability platforms such as Datadog, Prometheus, and Grafana.
- Experience supporting production systems in a 24/7/365 environment.
- Preferred experience with Kong Gateway, Kong Mesh, AWS networking, Azure VNet, GCP NCC, ClickHouse, Druid, PostgreSQL, Redis, disaster recovery, resiliency testing, and compliance-driven reliability practices.
Benefits
- Global 24/7 on-call rotation is part of the role.
- The position supports production SaaS systems across multiple regions and cloud providers.
Tech Stack
Apache KafkaAWSAzureBashClickHouseDatadogGoGoogle Cloud PlatformGrafanaHelmKubernetesLinuxPostgreSQLPrometheusPythonRedisTerraform
Categories
Site Reliability
