6 months ago
Remote, TurkeySenior
Responsibilities
- Own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
- Build and maintain CI/CD pipelines and deployment infrastructure.
- Use AI-driven automation to analyze and resolve production issues and improve development speed, reliability, and maintainability.
- Build dashboards, alerting, and anomaly detection across systems.
- Define and enforce SLOs and develop incident response processes.
- Manage and improve networking, load balancing, and service mesh configurations.
- Drive reliability improvements through automation, runbooks, and chaos engineering.
Requirements
- 5+ years of experience managing critical production systems and software development workflows.
- Strong production experience setting up and operating Kubernetes at scale with Terraform and Ansible.
- Deep knowledge of Linux networking, container networking including CNI plugins, VXLAN, and BGP, and DNS.
- Experience building CI/CD systems and GitOps workflows with FluxCD and ArgoCD.
- Proficiency in Python and either Go or Bash for tooling and automation.
- Strong experience with logging, monitoring, and alerting using Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, or Datadog.
- Excellent communication skills and ability to drive technical decisions across teams.
- Self-starter who executes quickly, takes ownership, and seeks continuous improvement.
- Nice to have experience managing GPU and AI/ML workloads.
- Nice to have experience with kernel-based monitoring and routing using eBPF and XDP.
- Nice to have experience with security tooling such as Falco, Coroot, and SIEM.
- Nice to have experience with bare-metal Kubernetes networking using Calico, Cilium, and MetalLB.
- Nice to have experience with distributed storage systems such as Ceph and Longhorn.
Benefits
- Work location in Turkey
- Learning and growth opportunities
- Regular team events and offsites
Tech Stack
Categories
DevOpsSite Reliability
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
