8 months ago
Remote, WorldwideSenior
Responsibilities
- Build and ship production agent capabilities for planning, tool use, memory, and context management.
- Integrate agents with retrieval systems, structured datasets, laboratory and biomedical APIs, spreadsheets, search, and other internal and external data sources.
- Develop unit, regression, scenario, benchmark, telemetry, and automated scoring systems for agent quality and evaluation.
- Collaborate with scientists to analyze failure modes and improve agent performance.
- Partner with knowledge and ontology teams to ensure source traceability and provenance compliance.
- Implement safety measures, guardrails, and sandboxed execution for risky operations.
- Improve performance and reliability through profiling, idempotency, retries, rate limiting, and uptime management.
- Instrument data pipelines for supervised fine-tuning and reinforcement learning when needed.
- Contribute to agent platform services, APIs, orchestration, CI/CD, and observability.
- Deliver multi-tool scientific agents, citation enforcement, and evaluation dashboards tracking competency, latency, and failure modes.
Requirements
- Production software development experience with Python and/or TypeScript and strong systems and API design skills.
- Experience shipping LLM applications or agentic systems involving tool use, retrieval/RAG, structured outputs, evaluation, or observability.
- Familiarity with agent and orchestration frameworks such as LangChain, LangGraph, AutoGen, CrewAI, and MCP, plus vector databases such as FAISS, Weaviate, and Pinecone.
- Experience with cloud infrastructure and containers, including AWS, GCP, or Azure, Docker, Kubernetes, Terraform, CI/CD, and production telemetry.
- Ability to translate research prototypes into robust, scalable systems.
- Experience with fine-tuning and reinforcement learning, including RL, RLAIF, RLHF, reward design, and offline evaluation, is preferred.
- Familiarity with SWE-Bench, OS-World, or tau-bench is preferred.
- Knowledge of retrieval and knowledge systems, schema and ontology design, entity modeling, and provenance tracking is preferred.
- Background in agentic system safety and security, including sandboxing, isolation, permissions, and auditability, is preferred.
- Exposure to life sciences or scientific computing and collaboration with domain experts is preferred.
Tech Stack
AWSAzureDockerFastAPIGoogle Cloud PlatformGraphQLgRPCKubernetesPostgreSQLPythonRedisTerraformTypeScript
