5 months ago
Responsibilities
- Own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
- Build and maintain CI/CD pipelines and deployment infrastructure.
- Use AI to automate production issue analysis and resolution and improve development speed, reliability, and maintainability.
- Build dashboards, alerting, and anomaly detection across systems.
- Define and enforce SLOs and develop incident response processes.
- Manage and improve networking, load balancing, and service mesh configurations.
- Drive reliability improvements through automation, runbooks, and chaos engineering.
Requirements
- 5+ years of experience managing critical production systems and software development workflows.
- Strong production experience setting up and operating Kubernetes at scale with Terraform and Ansible.
- Deep knowledge of Linux networking, container networking including CNI plugins, VXLAN, and BGP, and DNS.
- Experience building CI/CD systems and GitOps workflows using FluxCD and ArgoCD.
- Proficiency in Python and either Go or Bash for tooling and automation.
- Strong experience with logging, monitoring, and alerting using Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, and Datadog.
- Excellent communication skills and ability to drive technical decisions across teams.
- Preferred experience managing GPU and AI/ML workloads.
- Preferred experience with eBPF, XDP, Falco, Coroot, SIEM, Calico, Cilium, MetalLB, Ceph, or Longhorn.
Benefits
- Interesting and challenging work
- Learning and growth opportunities
- Regular team events and offsites
- Location: Turkey
