5 hours ago
Base Salary
$180k - $230k/yr
Responsibilities
- Extend state-of-the-art image, video, audio, and 3D models with new capabilities and modalities.
- Design model conditioning, generation, editing, training-free extensions, and fine-tuning approaches.
- Build and maintain fine-tuning APIs for customer model customization.
- Develop reusable abstractions and components for model capabilities and inference pipelines.
- Optimize generative model inference for low latency, high throughput, and efficient GPU utilization.
- Build, deploy, and maintain scalable and reliable generative model APIs.
- Collaborate directly with customers to develop generative media solutions.
- Take ambitious machine learning ideas independently from concept through production.
Requirements
- At least 3 years of professional experience as an Applied Machine Learning Engineer, including 1–2 years focused on generative media or computer vision.
- Expert-level proficiency in Python and PyTorch.
- Deep practical understanding of diffusion and flow-based generative models.
- Hands-on experience with open-weight model ecosystems, including Hugging Face and Diffusers.
- Experience deploying scalable, reliable, secure, safe, and performant machine learning systems to production.
- Experience designing training-free extensions for image, video, audio, or 3D generative models.
- Experience developing custom post-training or fine-tuning approaches for generative models.
- Experience in fast-paced startup environments or digital media and entertainment industries is desirable.
- Demonstrated ability to independently take ambitious machine learning ideas from concept to production.
Benefits
- Health, dental, and vision insurance in the US.
- Regular team events and offsites.
- Interesting and challenging work with learning and growth opportunities.
Categories
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
