1 month ago
Remote, WorldwideMid Level
Responsibilities
- Partner with the GM and early customers to define strong evaluations across domains.
- Work with researchers to design and build benchmarks and establish processing standards for different modalities.
- Build backend infrastructure including data pipelines, execution environments, storage, and orchestration.
- Create sandboxed environments for agentic evaluations involving tools, code execution, and multi-step tasks.
- Identify evaluation patterns, infrastructure gaps, and product opportunities from live customer engagements.
- Partner with DataLab on domain-specific data and research questions.
- Own the engineering portion of customer engagements end to end.
Requirements
- At least 4 years of engineering experience.
- Hands-on experience evaluating machine learning models.
- Previous ownership of backend and infrastructure.
- Ability to work effectively in high-ambiguity, urgent, fast-moving environments.
- Strong written communication skills.
- Preferred: experience building benchmarks, evaluations, or human data pipelines for LLMs.
- Preferred: experience at a frontier lab, evaluation-focused team, or research organization.
- Preferred: founding or early-engineer experience at a fast-moving startup.
- Preferred: familiarity with agentic systems, reinforcement-learning environments, code-execution sandboxes, and TEE/TREs.
Categories
Forward Deployed
