Snorkel AI

Senior / Staff AI Engineer

Snorkel AI
Apply
3 hours ago
San Francisco, CA, USA or New York, NY, USASenior / Staff+
H1B sponsor

Responsibilities

  • Design and build infrastructure for large-scale agentic workloads interacting with tools, external services, sandboxes, and simulated environments.
  • Build scalable synthetic data generation, automated labeling, dataset refinement, and evaluation systems.
  • Develop evaluation infrastructure for reproducible experiments, benchmark execution, regression detection, and continuous evaluation.
  • Build orchestration and distributed compute systems for thousands to millions of AI experiments and simulations across heterogeneous environments.
  • Develop agent simulation environments with provisioning, isolation, lifecycle management, and scalable execution.
  • Build and operate LLM infrastructure for routing, rate limiting, retries, caching, provider failover, cost attribution, and efficient multi-provider execution.
  • Instrument AI workloads with traces covering model interactions, tool calls, environment state, evaluation results, latency, reliability, and cost.
  • Improve developer experience through APIs, SDKs, workflow abstractions, and tooling for moving workloads from local development to production.
  • Collaborate with research, product, and engineering teams to turn experimental AI workflows into reliable platform capabilities.

Requirements

  • 5+ years building production software systems, including experience with AI/ML infrastructure, ML platforms, distributed systems, data platforms, or backend infrastructure.
  • Experience operating non-deterministic AI or ML workloads in production or at significant scale.
  • Experience building infrastructure for experimentation, evaluation, model development, synthetic data, agentic workflows, training, inference, or production ML systems.
  • Strong proficiency in Python and experience building production-quality APIs, services, and developer tooling.
  • Strong background in distributed systems and cloud platforms, including compute orchestration, storage, networking, isolation, and failure handling; AWS is preferred.
  • Experience with workflow or distributed execution frameworks such as Prefect, Airflow, Dagster, Ray, Kubernetes, or similar systems.
  • Strong understanding of observability, telemetry, reliability, performance, debugging, incident response, and cost management.
  • Ability to evaluate AI system quality through evaluation design, experiment reproducibility, behavioral regression detection, and analysis of model or agent variability.
  • Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact.
  • Strong technical communication skills and ability to work in a fast-paced environment.
  • Fluency with modern AI and developer tooling and willingness to evaluate and adopt evolving models, frameworks, infrastructure, and techniques.
  • Nice-to-have experience with LLM or agent infrastructure, evaluation platforms, synthetic data, reinforcement learning environments, large-scale distributed AI workloads, LLM observability, sandboxing, shared AI platform libraries or SDKs, hyper-growth startups, or prior Tech Lead, Team Lead, or hands-on Engineering Manager experience.

Benefits

  • Meaningful ownership over foundational AI infrastructure and architecture decisions.
  • Opportunities to shape priorities, influence strategic decisions, deepen technical expertise, and pursue leadership or learning opportunities.
  • The company offers a scaling environment with market-proven solutions, robust funding, and high-growth opportunities.

Tech Stack

Apache AirflowAWSKubernetesPython
Snorkel AI

About Snorkel AI

1,001-5,000 employees

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine platform technology with research-driven data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. Snorkel led the development of Senior SWE-Bench and launched Open Benchmarks Grants with a $3 million commitment to support open-source datasets, benchmarks, and evaluation research. Supported projects include Agents’ Last Exam, OSWorld 2.0, Terminal-Bench, Continual Learning Bench, and SlopCode Bench. Learn more at snorkel.ai or follow @SnorkelAI.

Contact me