1 month ago
Remote, United States or San Francisco, CA, USASenior
Responsibilities
- Operate EKS clusters across environments using Karpenter autoscaling, Cilium networking, and ArgoCD-driven deployments.
- Manage a multi-account AWS organization, including provisioning, networking, access control, and cross-account connectivity.
- Maintain and evolve Terraform/Terragrunt infrastructure-as-code modules, state management, and plan/apply pipelines.
- Improve tooling for deployments, schema changes, backups, restores, and incident response.
- Identify recurring operational problems and eliminate them through automation and self-healing systems.
- Optimize cloud spending and participate in on-call and incident response while reducing incident frequency over time.
Requirements
- Deep hands-on production Kubernetes experience, preferably with EKS, including debugging node pressure, networking issues, and deployment failures at scale.
- Strong experience operating production infrastructure on AWS across multiple accounts, including organizational boundaries, IAM, and networking.
- Experience automating infrastructure with Terraform or Terragrunt at scale, including module design and state management.
- Solid understanding of Linux systems, including disk, memory, networking, and failure modes.
- Experience supporting stateful systems such as databases, queues, or storage systems.
- Ability to debug and reason about production performance and reliability issues and own systems end-to-end, including on-call responsibilities.
- Nice-to-have experience with ArgoCD, GitHub Actions, multi-region infrastructure, consistency and availability tradeoffs, or AI agent-enabled infrastructure services.
Benefits
- Natively remote position for candidates in the US Pacific timezone.
- Async-first work culture with Tuesdays and Thursdays meeting-free and an emphasis on heads-down building time.
- Autonomous, product-led environment where engineers own projects and make product decisions.
