Responsibilities
- Design scalable, highly available AI platform infrastructure covering compute, storage, networking, security, enterprise integration patterns, Terraform and Helm, multi-region high availability/disaster recovery, and CI/CD
- Design and implement multi-agent systems, agent logic, evaluation frameworks, prompt optimization, and deployment and operations practices
- Lead technical maturity assessments, understand enterprise customer requirements, and present recommendations
- Partner with Engagement Managers, Product teams, and Engineering teams to guide customer implementations
- Balance infrastructure architecture with agent development to solve customer business problems
Requirements
- At least 7 years of experience in hands-on technical customer-facing roles such as Solutions Architect or Forward Deployed Engineer
- At least 3 years of experience designing and deploying production infrastructure on GCP, AWS, or Azure
- Strong Kubernetes experience, including cluster design, autoscaling, and multi-zone deployments
- Experience with Terraform, Helm, GitOps practices, high availability, disaster recovery, networking, security, observability, and CI/CD pipelines
- Knowledge of relational databases and in-memory data stores, including replication, backup strategies, and sizing
- At least 1 year of experience building production AI/ML applications or agents
- Strong experience with LangChain, LangGraph, or similar LLM frameworks
- Experience with state management, AI evaluation frameworks, prompt optimization, A/B testing, vector stores, RAG patterns, tool integration, API design, and error handling
- Strong Python and/or TypeScript development skills
- Experience working with enterprise customers, conducting technical assessments or infrastructure audits, and explaining technical concepts to diverse audiences
- Strong problem-solving, communication, cross-functional collaboration, consultative, and hands-on engineering skills
Benefits
- Medical, dental, and vision coverage
- Flexible vacation
- 401(k) plan
- Meals on in-office days in the US
- Remote work with up to 25% travel
- Locally competitive benefits for team members in the EU, UK, and APAC
Tech Stack
Categories
About LangChain
At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. What began as widely adopted open-source tools has grown into a platform for building, evaluating, deploying, and operating agents at scale. LangChain provides the agent engineering platform and open source frameworks developers need to ship reliable agents fast. LangSmith offers observability, evaluation, and deployment for rapid iteration. Our open source frameworks, LangGraph, LangChain, and Deep Agents, help developers build agents with speed and granular control. LangSmith is trusted by leading AI teams at Zip, Vanta, Klarna, Workday, Linkedin, Cloudflare, and more.
