Snorkel AI

Senior/Staff FDE - Synthetic Data Generation

Snorkel AI
Apply
4 hours ago
San Francisco, CA, USA or New York, NY, USASenior / Staff+
H1B Sponsor

Base Salary

$180k - $320k/yr

Responsibilities

  • Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines.
  • Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications.
  • Develop LLM- and ML-assisted workflows for generating training and evaluation datasets.
  • Build automated evaluators, quality checks, and measurement frameworks for correctness, relevance, diversity, coverage, and customer requirements.
  • Run experiments to measure synthetic data impact on downstream model performance and improve generation approaches.
  • Deliver production-grade datasets with standardized formats, quality assurance, and documentation.
  • Lead customer technical workstreams from solution design through production delivery.
  • Prototype and productionize solutions across models, data pipelines, APIs, and custom applications.
  • Communicate technical tradeoffs, experimental results, and recommendations to technical and cross-functional stakeholders.
  • Turn recurring customer solutions into reusable pipelines, evaluators, tooling, standards, and best practices.
  • Lead technical design reviews, guide other engineers, and influence product and platform capabilities.

Requirements

  • 5+ years of experience in machine learning engineering, data science, applied AI, forward deployed engineering, or a similar technical role.
  • Strong Python skills and experience building reliable production data or ML systems.
  • Experience containerizing systems with Docker and deploying them on AWS, GCP, or Azure.
  • Hands-on experience building LLM-based applications and data workflows and integrating systems, models, and data sources through APIs.
  • Strong understanding of ML experimentation and evaluation, including defining metrics and using empirical results to guide decisions.
  • Experience building synthetic data, data augmentation, or model-generated training and evaluation datasets.
  • Experience with LLM evaluation techniques such as LLM-as-a-judge, model-based evaluation, rubric-based evaluation, or custom evaluators.
  • Ability to take ambiguous technical problems from definition through delivery and work directly with customers and cross-functional stakeholders.
  • Experience setting technical direction, driving architecture and key decisions, mentoring engineers, and creating reusable approaches.
  • Preferred: experience developing fine-tuning, preference optimization, or benchmarking datasets with human-in-the-loop workflows.
  • Preferred: experience building agentic environments and tasks, including repo-scale coding tasks, tool-agent-user interaction design, and agent tool protocols.
  • Preferred: experience with reinforcement learning for LLMs, reward and verifier design, or RL with verifiable rewards.
  • Preferred: experience in fast-paced customer-facing environments with evolving requirements and technical approaches.

Benefits

  • Base salary range of $180,000–$320,000, plus variable compensation opportunity; compensation details are excluded from benefits classification.
  • Employee stock options and benefits are included with offers.
  • Opportunities to shape priorities, influence strategic decisions, and impact company initiatives.
  • Support for deepening technical expertise, exploring leadership opportunities, and learning across functions.
  • Equal employment opportunity and reasonable accommodation are provided.

Categories

Forward Deployed
Snorkel AI

About Snorkel AI

201-500 employees

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine platform technology with research-driven data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. Snorkel led the development of Senior SWE-Bench and launched Open Benchmarks Grants with a $3 million commitment to support open-source datasets, benchmarks, and evaluation research. Supported projects include Agents’ Last Exam, OSWorld 2.0, Terminal-Bench, Continual Learning Bench, and SlopCode Bench. Learn more at snorkel.ai or follow @SnorkelAI.

Contact me