fal

Software Engineer, Distributed Systems

fal
Apply
6 months ago
Remote, TurkeySenior

Responsibilities

  • Build the core Python and Rust platform for request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing.
  • Design the platform’s evolution to support 100x current traffic and low latency worldwide.
  • Use AI extensively to automate mundane aspects of building complex, reliable systems.
  • Profile and tune low-level CPU and memory performance.

Requirements

  • 5+ years of experience building distributed compute and orchestration platforms in Python or Rust.
  • Strong understanding of distributed systems fundamentals, including consensus, scheduling, fault tolerance, and capacity planning.
  • Deep understanding of computational complexity and memory allocation.
  • Track record of designing systems that scale under real production load.
  • Experience building and using observability to guide performance and reliability decisions.
  • Excellent communication skills and ability to drive technical decisions across teams.
  • Self-starter who executes quickly, takes ownership, and seeks continuous improvement.
  • Experience with AI/ML inference or training infrastructure is a nice-to-have.
  • Experience with high-performance systems programming, including async runtimes, zero-copy, or memory-safe concurrency, is a nice-to-have.
  • Background building multi-tenant compute platforms is a nice-to-have.
  • Understanding of networking fundamentals and performance characteristics is a nice-to-have.
  • Familiarity with GPU workload characteristics and scheduling constraints is a nice-to-have.

Benefits

  • Location: Turkey.
  • Interesting and challenging work.
  • Learning and growth opportunities.
  • Regular team events and offsites.

Tech Stack

Categories

fal

About fal

51-200 employees

Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.

Contact me