Featherless AI

Machine Learning Engineer — Multilingual Data

Featherless AI
Apply
8 months ago
Remote, United StatesMid Level

Responsibilities

  • Design, build, and maintain large-scale multilingual datasets across high- and low-resource languages.
  • Develop pipelines for data collection, cleaning, normalization, deduplication, and labeling.
  • Implement statistical, heuristic, and model-based quality filters.
  • Define language coverage, benchmarks, and evaluation metrics with researchers.
  • Analyze dataset bias, coverage gaps, and failure modes across regions and scripts.
  • Support model training, fine-tuning, and distillation workflows with high-quality multilingual data.
  • Iterate on datasets based on model performance and real-world usage.

Requirements

  • At least 3 years of experience as an ML Engineer, Applied Scientist, or in a similar role.
  • Strong experience working with multilingual or non-English datasets.
  • Solid understanding of NLP fundamentals including tokenization, embeddings, and language modeling.
  • Experience building scalable data pipelines with Python, Spark, Ray, or similar technologies.
  • Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks.
  • Ability to collaborate with researchers and translate research needs into production systems.
  • Experience with low-resource languages or multilingual benchmarks such as FLORES or XTREME is preferred.
  • Exposure to LLM training, fine-tuning, or distillation is preferred.
  • A linguistics background or experience working with native language experts is preferred.
  • Contributions to open-source datasets or ML tooling and experience with data quality evaluation at scale are preferred.

Benefits

  • Real ownership over a core product differentiator.
  • Work on models used globally across English and non-English markets.
  • Small, high-caliber team with deep ML and systems experience.
  • Competitive compensation and meaningful equity at a Series A stage company.

Categories

Data EngineeringML Engineering
Featherless AI

About Featherless AI

11-50 employees

Featherless AI builds a serverless inference platform that orchestrates GPUs and load balances models so teams can deploy and scale open‑source AI without managing infrastructure. Its public cloud serves tens of thousands of open‑weight models and supports fine‑tuning, targeting developers, ML engineers, and enterprises needing reliable, high‑throughput inference. Founded in 2023 and headquartered in San Francisco, the privately held, Series A company is backed by investors including AMD and Airbus Ventures.

Contact me