6 months ago
Base Salary
$180k - $250k/yr
Responsibilities
- Own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
- Build and maintain CI/CD pipelines and deployment infrastructure.
- Use AI to automate production issue analysis and resolution while improving development speed, reliability, and maintainability.
- Build dashboards, alerting, and anomaly detection across systems.
- Define and enforce SLOs and establish incident response processes.
- Manage and improve networking, load balancing, and service mesh configurations.
- Drive reliability improvements through automation, runbooks, and chaos engineering.
Requirements
- At least five years of experience managing critical production systems and software development workflows.
- Strong production experience setting up and operating Kubernetes at scale with Terraform and Ansible.
- Deep knowledge of Linux networking, container networking including CNI plugins, VXLAN, and BGP, and DNS.
- Experience building CI/CD systems and GitOps workflows using FluxCD and ArgoCD.
- Proficiency in Python and either Go or Bash for tooling and automation.
- Strong experience with logging, monitoring, and alerting using Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, or Datadog.
- Excellent communication skills and the ability to drive technical decisions across teams.
- Nice-to-have experience managing GPU and AI/ML workloads, kernel-based monitoring and routing with eBPF or XDP, security tooling such as Falco, Coroot, or SIEM, bare-metal Kubernetes networking with Calico, Cilium, or MetalLB, and distributed storage systems such as Ceph or Longhorn.
Benefits
- Interesting and challenging work with learning and growth opportunities.
- Health, dental, and vision insurance in the US.
- Regular team events and offsites.
- Role based in downtown San Francisco, California.
