Snorkel AI

Senior/Staff FDE - CUA

Snorkel AI
Apply
2 hours ago
San Francisco, CA, USA or New York, NY, USASenior / Staff+
H1B Sponsor

Base Salary

$180k - $320k/yr

Responsibilities

  • Design and build task environments, datasets, evaluation workflows, automated evaluators, and measurement frameworks for computer-use agents.
  • Translate customer goals and agent failure modes into representative multi-step tasks with clear success criteria.
  • Develop data-generation, validation, and quality-assurance pipelines for multimodal and agentic training and evaluation data.
  • Diagnose failures involving planning, tool use, perception, state management, and user-interface interaction, then improve tasks, data, and evaluations.
  • Lead technical workstreams from discovery and solution design through implementation, evaluation, and production delivery.
  • Prototype and productionize solutions across models, agent frameworks, APIs, browser or desktop environments, and custom applications.
  • Communicate technical tradeoffs and experimental results to customers and cross-functional stakeholders.
  • Create reusable task frameworks, evaluators, tooling, technical standards, and best practices across engagements.
  • Lead technical design reviews, guide other engineers, and influence platform and product capabilities.

Requirements

  • 5+ years of experience in machine learning engineering, software engineering, applied AI, forward deployed engineering, solutions engineering, or a similar technical role.
  • Strong Python skills and experience building reliable production software, data, or ML systems.
  • Hands-on experience building, evaluating, or deploying LLM-based or agentic systems, including computer-use agents.
  • Strong understanding of experimentation and evaluation, including LLM-as-a-judge or model-based evaluation, metrics, and empirical decision-making.
  • Experience designing agent task environments, datasets, verifiers, and reward or verifier systems.
  • Experience with APIs, automation, web applications, browser-based systems, desktop applications, or developer tools.
  • Experience manipulating, analyzing, and validating large or complex datasets using Python and the modern GenAI/LLM stack.
  • Ability to take ambiguous technical problems from definition through implementation and delivery.
  • Strong technical communication skills and experience working directly with customers or cross-functional stakeholders.
  • Experience setting technical direction, creating reusable approaches, and influencing engineering or product decisions.
  • Preferred: experience with agent benchmarks, task suites, simulators, multimodal models, visual grounding, graphical-user-interface agents, repo-scale coding tasks, agent tool protocols, browser/computer-use products, and data pipelines for fine-tuning, reinforcement learning, preference optimization, benchmarking, or model evaluation.

Benefits

  • Base compensation includes an additional variable compensation opportunity, equity in the form of employee stock options, and benefits.
  • The company offers opportunities for career growth, learning, increased technical expertise, and leadership development.
  • Snorkel AI is an equal employment opportunity employer and provides reasonable accommodations for individuals with disabilities.

Tech Stack

Categories

Forward Deployed
Snorkel AI

About Snorkel AI

201-500 employees

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine platform technology with research-driven data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. Snorkel led the development of Senior SWE-Bench and launched Open Benchmarks Grants with a $3 million commitment to support open-source datasets, benchmarks, and evaluation research. Supported projects include Agents’ Last Exam, OSWorld 2.0, Terminal-Bench, Continual Learning Bench, and SlopCode Bench. Learn more at snorkel.ai or follow @SnorkelAI.

Contact me