Kong

Site Reliability Engineer 2

Kong
Apply
10 days ago
Washington, DC, USAMid Level

Base Salary

$123k - $150k/yr

Responsibilities

  • Operate and scale Kong’s multi-region, multi-tenant Konnect SaaS platform across AWS, GCP, and Azure.
  • Build and maintain Kubernetes-based infrastructure and deployment workflows using Terraform, Terragrunt, Helm, and ArgoCD.
  • Design and optimize multi-region PostgreSQL, Redis, ClickHouse, and Druid data and caching layers.
  • Operate Kong Gateway and Kong Mesh environments supporting hybrid and distributed architectures.
  • Develop CI/CD pipelines and GitOps workflows for automated service delivery and infrastructure changes.
  • Improve observability and incident response using Datadog, Prometheus, Grafana, and Thanos while defining and tracking SLOs.
  • Collaborate with development and security teams on reliable, secure, and compliant SaaS operations.
  • Participate in a global 24/7 on-call rotation and improve operational playbooks and postmortem practices.
  • Lead scaling initiatives that improve platform elasticity, reliability, resilience, and cost efficiency.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • Experience managing enterprise-scale SaaS or PaaS systems in multi-region, multi-tenant, secure environments.
  • Deep Kubernetes expertise, including cluster and networking troubleshooting and fault-tolerant, scalable design.
  • Strong proficiency with Terraform or Terragrunt and experience with CI/CD and GitOps workflows using ArgoCD, Atlantis, and Helm.
  • Proficiency in one or more of Go, Python, or Bash for automation and tooling.
  • Understanding of Linux/Unix, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.
  • Experience with API gateway and service mesh technologies.
  • Familiarity with Kafka and observability platforms such as Datadog, Prometheus, and Grafana.
  • Experience supporting production systems in a 24/7/365 environment.
  • Preferred experience with Kong Gateway, Kong Mesh, AWS networking, Azure VNet, GCP NCC, ClickHouse, Druid, PostgreSQL, Redis, disaster recovery, resiliency testing, and compliance-driven reliability practices.

Benefits

  • Global 24/7 on-call rotation is part of the role.
  • The position supports production SaaS systems across multiple regions and cloud providers.

Tech Stack

Categories

Site Reliability
Kong

About Kong

1,001-5,000 employees
Contact me