Base Salary
$220k - $320k/yr
Responsibilities
- Lead projects from data intake through the full model-training pipeline, including dataset processing, cleaning, and preparation.
- Build and maintain pipelines for aggregating, transforming, and validating training data.
- Create dashboards and visualization tools for training metrics, data quality, and model performance.
- Train models with internal frameworks and iterate based on evaluation results.
- Develop benchmarks and evaluation frameworks to measure frontier-level model performance.
- Automate portions of the training workflow to reduce manual intervention and improve consistency.
- Take research features into production settings.
- Apply SFT, RL, and model-optimization techniques to improve training quality and efficiency.
- Collaborate with infrastructure engineers to scale training across the GPU fleet.
- Understand customer use cases to guide training strategies and identify edge cases.
Requirements
- At least 2 years of experience training AI models using PyTorch.
- Hands-on experience post-training LLMs using supervised fine-tuning or reinforcement learning.
- Strong understanding of transformer architectures and how they are trained.
- Experience with LLM training frameworks such as Hugging Face Transformers, DeepSpeed, or Axolotl.
- Experience training models on NVIDIA GPUs.
- Strong data-processing skills, including building ETL pipelines and working with large datasets.
- A track record of creating benchmarks and evaluations.
- Ability to apply research techniques to production systems.
- Preferred: experience with model distillation or knowledge transfer.
- Preferred: experience building dashboards and data-visualization tools.
- Preferred: familiarity with vision encoders and multimodal models.
- Preferred: experience with distributed training at scale.
- Preferred: contributions to open-source ML projects.
Benefits
- Equity and comprehensive benefits are offered.
- The team works in person in downtown San Francisco, typically four days per week; hybrid work is available for Bay Area candidates.
Tech Stack
Categories
About Inference
Inference.net helps teams ship AI that’s faster, smarter, and dramatically more cost-efficient. We deliver lower latency and higher-quality models at a fraction of the cost, with full OpenAI compatibility and no vendor lock-in. Companies use Inference.net to power real-time AI features, automate workflows, and scale mission-critical systems without blowing up their margins. Case studies: • The fastest growing AEO optimization software: trained a custom model that improved ranking accuracy for Fortune 500 clients, increasing conversion while lowering inference spend and stabilizing latency across high-volume workloads. • Fastest-growing nutrition tracking app: scaled to 10M+ users while cutting inference costs, processing millions of daily food images with custom vision models outperforming frontier VLMs. • Decentralized data network (search.video): processed billions of monthly video frames across 3M+ nodes using a specialized ClipTagger that boosted relevance and throughput. • Digital bank (120M+ customers): delivered 99.99%+ uptime, faster and more accurate compliance/servicing models at a fraction of API cost. Build AI products that scale, without sacrificing performance or profitability. You can find us also on X: https://x.com/inference_net
