Featherless AI

Machine Learning Engineer — Distillation

Featherless AI
Apply
8 months ago
Remote, WorldwideMid Level

Responsibilities

  • Design and implement teacher-student, self-distillation, and multi-teacher knowledge distillation pipelines.
  • Distill large foundation models into smaller, faster, and cheaper inference models.
  • Run and analyze large-scale training experiments to evaluate quality, latency, and cost tradeoffs.
  • Collaborate with research teams to translate new distillation ideas into production-ready code.
  • Optimize training and inference performance across memory, throughput, and latency.
  • Contribute to internal tooling, evaluation frameworks, and experiment tracking.
  • Optionally contribute to open-source models, tooling, or research.

Requirements

  • Strong background in machine learning or deep learning.
  • Hands-on experience with model distillation for LLMs or other neural networks.
  • Solid understanding of training dynamics, loss functions, and optimization.
  • Experience with PyTorch or JAX and modern machine learning tooling.
  • Comfort running experiments on multi-GPU or distributed setups.
  • Ability to reason about model quality versus performance tradeoffs.
  • Experience distilling LLMs or large sequence models is preferred.
  • Experience with inference optimization, including quantization, pruning, or kernels, is preferred.
  • Familiarity with language-model evaluation is preferred.
  • Open-source contributions or research publications are preferred.
  • Experience in early-stage or fast-moving startups is preferred.

Benefits

  • Competitive compensation and meaningful equity.
  • Remote-friendly, async-first environment.
  • High ownership and direct impact on product and roadmap.
  • Opportunity to work with a small, senior team in a research-and-engineering culture.

Tech Stack

Categories

Featherless AI

About Featherless AI

11-50 employees

Featherless AI builds a serverless inference platform that orchestrates GPUs and load balances models so teams can deploy and scale open‑source AI without managing infrastructure. Its public cloud serves tens of thousands of open‑weight models and supports fine‑tuning, targeting developers, ML engineers, and enterprises needing reliable, high‑throughput inference. Founded in 2023 and headquartered in San Francisco, the privately held, Series A company is backed by investors including AMD and Airbus Ventures.

Contact me