
Featherless AI
Featherless AI builds a serverless inference platform that orchestrates GPUs and load balances models so teams can deploy and scale open‑source AI without managing infrastructure. Its public cloud serves tens of thousands of open‑weight models and supports fine‑tuning, targeting developers, ML engineers, and enterprises needing reliable, high‑throughput inference. Founded in 2023 and headquartered in San Francisco, the privately held, Series A company is backed by investors including AMD and Airbus Ventures.
Open Positions at Featherless AI
5 open positions
Research and build next-generation neural network architectures, moving ideas from experiments into scalable production systems. This role is suited to an engineer who wants to work beyond standard Transformers across theory, model implementation, and deployment.
Build and ship smaller, faster, and more efficient models through hands-on knowledge distillation work. You’ll bridge research and production by designing pipelines, running large-scale experiments, and optimizing models for real-world inference.
Own the performance of large-scale ML inference systems, turning advanced models into fast, reliable, and cost-efficient production services. This hands-on role spans GPU-level profiling, inference optimization, model serving, and productionizing research breakthroughs.
Optimize large-scale model training systems for speed, stability, and cost while working across research and production. You’ll own distributed training performance and build robust infrastructure for scaling advanced models.
Own and scale multilingual data pipelines that improve model quality across languages, scripts, and cultural contexts. You’ll bridge data, research, and production ML while working with Python, Spark, or Ray.