16 days ago
Responsibilities
- Design and run experiments across agent behavior, traces, and activations to identify failures and develop interventions that improve intelligence, cost, and reliability.
- Research context and memory management, agent steering, task decomposition and specialization, continual learning during deployment, and inter-agent communication.
- Build durable agent analysis and evaluation infrastructure, including observability, behavior analysis, evaluation harnesses, and auditing tools.
- Publish evaluation results, methods, and discoveries as company assets and public research.
- Collaborate across teams to design scalable, cost-efficient, low-latency systems and novel model or agent architectures for heterogeneous hardware.
Requirements
- Demonstrated ability to independently conduct research by formulating open questions, designing experiments, and producing reusable results.
- Deep hands-on experience with LLMs in agentic settings, including multi-step tasks, tool use, and long-horizon workflows.
- Strong experimental discipline with hypotheses, controlled comparisons, ablations, and careful treatment of uncertainty and benchmark noise.
- Strong software engineering skills and comfort writing experiment code within a substantial shared codebase.
- Strong communication skills for translating research findings into guidance and communicating them to wider audiences.
Benefits
- Competitive salary determined by skills and experience, equity and ownership, and private healthcare.
- Visa sponsorship and relocation benefits are offered.
- The role is based in person at the company's London office with provided tools, workspace, and setup.
Categories
AI Research
