Inference

Applied Machine Learning Engineer

Inference
Apply
7 months ago
San Francisco, CA, USAMid Level
H1B Sponsor

Base Salary

$220k - $320k/yr

Responsibilities

  • Lead projects from data intake through the full model-training pipeline, including dataset processing, cleaning, and preparation.
  • Build and maintain pipelines for aggregating, transforming, and validating training data.
  • Create dashboards and visualization tools for training metrics, data quality, and model performance.
  • Train models with internal frameworks and iterate based on evaluation results.
  • Develop benchmarks and evaluation frameworks to measure frontier-level model performance.
  • Automate portions of the training workflow to reduce manual intervention and improve consistency.
  • Take research features into production settings.
  • Apply SFT, RL, and model-optimization techniques to improve training quality and efficiency.
  • Collaborate with infrastructure engineers to scale training across the GPU fleet.
  • Understand customer use cases to guide training strategies and identify edge cases.

Requirements

  • At least 2 years of experience training AI models using PyTorch.
  • Hands-on experience post-training LLMs using supervised fine-tuning or reinforcement learning.
  • Strong understanding of transformer architectures and how they are trained.
  • Experience with LLM training frameworks such as Hugging Face Transformers, DeepSpeed, or Axolotl.
  • Experience training models on NVIDIA GPUs.
  • Strong data-processing skills, including building ETL pipelines and working with large datasets.
  • A track record of creating benchmarks and evaluations.
  • Ability to apply research techniques to production systems.
  • Preferred: experience with model distillation or knowledge transfer.
  • Preferred: experience building dashboards and data-visualization tools.
  • Preferred: familiarity with vision encoders and multimodal models.
  • Preferred: experience with distributed training at scale.
  • Preferred: contributions to open-source ML projects.

Benefits

  • Equity and comprehensive benefits are offered.
  • The team works in person in downtown San Francisco, typically four days per week; hybrid work is available for Bay Area candidates.

Tech Stack

Hugging Face TransformersPyTorch

Categories

Inference

About Inference

11-50 employees

Inference.net helps teams ship AI that’s faster, smarter, and dramatically more cost-efficient. We deliver lower latency and higher-quality models at a fraction of the cost, with full OpenAI compatibility and no vendor lock-in. Companies use Inference.net to power real-time AI features, automate workflows, and scale mission-critical systems without blowing up their margins. Case studies: • The fastest growing AEO optimization software: trained a custom model that improved ranking accuracy for Fortune 500 clients, increasing conversion while lowering inference spend and stabilizing latency across high-volume workloads. • Fastest-growing nutrition tracking app: scaled to 10M+ users while cutting inference costs, processing millions of daily food images with custom vision models outperforming frontier VLMs. • Decentralized data network (search.video): processed billions of monthly video frames across 3M+ nodes using a specialized ClipTagger that boosted relevance and throughput. • Digital bank (120M+ customers): delivered 99.99%+ uptime, faster and more accurate compliance/servicing models at a fraction of API cost. Build AI products that scale, without sacrificing performance or profitability. You can find us also on X: https://x.com/inference_net