1 year ago
Base Salary
$180k - $250k/yr
Responsibilities
- Build fal’s core Python and Rust platform for request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing.
- Create forward-looking designs for platform evolution as traffic scales by 100x while maintaining low latency worldwide.
- Use AI extensively to automate routine aspects of building complex and reliable systems.
- Profile and tune low-level CPU and memory performance.
Requirements
- At least 3 years of experience building distributed compute and orchestration platforms in Python or Rust.
- Strong understanding of distributed systems fundamentals, including consensus, scheduling, fault tolerance, and capacity planning.
- Deep understanding of computational complexity and memory allocation.
- Track record of designing systems that scale under real production load.
- Experience building and using observability to guide performance and reliability decisions.
- Excellent communication skills and ability to drive technical decisions across teams.
- Self-starter who executes quickly, takes ownership, and seeks continuous improvement.
- Experience with AI/ML inference or training infrastructure is preferred.
- Experience with high-performance systems programming, including async runtimes, zero-copy techniques, and memory-safe concurrency, is preferred.
- Background building multi-tenant compute platforms is preferred.
- Understanding of networking fundamentals and performance characteristics is preferred.
- Familiarity with GPU workload characteristics and scheduling constraints is preferred.
Benefits
- The role is based in downtown San Francisco, California.
- Relocation assistance to San Francisco is offered.
- Health, dental, and vision insurance are offered in the US.
- Regular team events and offsites are provided.
- The role offers challenging work and learning and growth opportunities.
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
