GrepJob
fal

Software Engineer, Site Reliability

fal
Apply
5 months ago
Remote, WorldwideSenior
H1B Sponsor

Responsibilities

  • Own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
  • Build and maintain CI/CD pipelines and deployment infrastructure.
  • Use AI to automate production issue analysis and resolution and improve development speed, reliability, and maintainability.
  • Build dashboards, alerting, and anomaly detection across systems.
  • Define and enforce SLOs and develop incident response processes.
  • Manage and improve networking, load balancing, and service mesh configurations.
  • Drive reliability improvements through automation, runbooks, and chaos engineering.

Requirements

  • 5+ years of experience managing critical production systems and software development workflows.
  • Strong production experience setting up and operating Kubernetes at scale with Terraform and Ansible.
  • Deep knowledge of Linux networking, container networking including CNI plugins, VXLAN, and BGP, and DNS.
  • Experience building CI/CD systems and GitOps workflows using FluxCD and ArgoCD.
  • Proficiency in Python and either Go or Bash for tooling and automation.
  • Strong experience with logging, monitoring, and alerting using Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, and Datadog.
  • Excellent communication skills and ability to drive technical decisions across teams.
  • Preferred experience managing GPU and AI/ML workloads.
  • Preferred experience with eBPF, XDP, Falco, Coroot, SIEM, Calico, Cilium, MetalLB, Ceph, or Longhorn.

Benefits

  • Interesting and challenging work
  • Learning and growth opportunities
  • Regular team events and offsites
  • Location: Turkey

Tech Stack

AnsibleBashDatadogGoGrafanaKubernetesLinuxPrometheusPythonTerraform

Categories