5 months ago
Responsibilities
- Build fal’s core Python and Rust platform for request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing.
- Design the platform’s evolution to support 100x current traffic while maintaining low latency globally.
- Use AI extensively to automate mundane aspects of building complex, reliable systems.
- Profile and tune low-level CPU and memory performance.
Requirements
- At least five years of experience building distributed compute and orchestration platforms in Python or Rust.
- Strong understanding of distributed systems fundamentals, including consensus, scheduling, fault tolerance, and capacity planning.
- Deep understanding of computational complexity and memory allocation.
- Track record of designing systems that scale under real production load.
- Experience building and using observability to guide performance and reliability decisions.
- Excellent communication skills and ability to drive technical decisions across teams.
- Self-starter who executes quickly, takes ownership, and seeks continuous improvement.
- Preferred: experience with AI/ML inference or training infrastructure.
- Preferred: experience with high-performance systems programming, including async runtimes, zero-copy, or memory-safe concurrency.
- Preferred: background building multi-tenant compute platforms.
- Preferred: understanding of networking fundamentals and performance characteristics.
- Preferred: familiarity with GPU workload characteristics and scheduling constraints.
Benefits
- Interesting and challenging work
- Learning and growth opportunities
- Regular team events and offsites
- Location: Turkey
