fal

Software Engineer, Site Reliability

fal
Apply
7 months ago

Base Salary

$180k - $250k/yr

Responsibilities

  • Own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
  • Build and maintain deployment infrastructure and CI/CD pipelines.
  • Use AI extensively to automate production issue analysis and resolution and improve software development speed, reliability, and maintainability.
  • Build dashboards, alerting, and anomaly detection across systems.
  • Define and enforce SLOs and develop incident response processes.
  • Manage and improve networking, load balancing, and service mesh configurations.
  • Drive reliability improvements through automation, runbooks, and chaos engineering.

Requirements

  • At least 5 years of experience managing critical production systems and software development workflows.
  • Strong production experience setting up and operating Kubernetes at scale with Terraform and Ansible.
  • Deep knowledge of Linux networking, container networking including CNI plugins, VXLAN, and BGP, and DNS.
  • Experience building CI/CD systems and GitOps workflows with FluxCD and ArgoCD.
  • Proficiency in Python and either Go or Bash for tooling and automation.
  • Strong experience with logging, monitoring, and alerting tools including Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, and Datadog.
  • Excellent communication skills and the ability to drive technical decisions across teams.
  • Preferred experience managing GPU and AI/ML workloads, using eBPF and XDP for kernel-based monitoring and routing, and working with Falco, Coroot, Calico, Cilium, MetalLB, Ceph, or Longhorn.

Benefits

  • Health, dental, and vision insurance for US employees.
  • Regular team events and offsites.
  • Hiring is currently in downtown San Francisco, California.
  • Offers interesting and challenging work with learning and growth opportunities.

Tech Stack

AnsibleBashDatadogGoGrafanaKubernetesLinuxPrometheusPythonTerraform

Categories

DevOpsSite Reliability
fal

About fal

51-200 employees

Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.

Contact me