about 1 year ago
Base Salary
$180k - $250k/yr
Responsibilities
- Build fal’s core Python/Rust platform for request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing
- Design the platform’s evolution to support 100 times current traffic with low latency worldwide
- Use AI extensively to automate routine aspects of building complex and reliable systems
- Profile and tune low-level CPU and memory performance
Requirements
- At least 3 years of experience building distributed compute and orchestration platforms in Python or Rust
- Strong understanding of consensus, scheduling, fault tolerance, and capacity planning
- Deep understanding of computational complexity and memory allocation
- Track record designing systems that scale under real production load
- Experience building and using observability to guide performance and reliability decisions
- Excellent communication skills and ability to drive technical decisions across teams
- Experience with AI/ML inference or training infrastructure is preferred
- Experience with high-performance systems programming, including async runtimes, zero-copy techniques, and memory-safe concurrency, is preferred
- Background building multi-tenant compute platforms is preferred
- Understanding of networking fundamentals and performance characteristics is preferred
- Familiarity with GPU workload characteristics and scheduling constraints is preferred
Benefits
- Health, dental, and vision insurance in the US
- Relocation assistance to San Francisco
- Regular team events and offsites
- On-site work in downtown San Francisco, California
