2 months ago
Durham, NC, USA or Sunnyvale, CA, USAStaff+
Base Salary
$147k - $224k/yr
Responsibilities
- Lead the design, development, deployment, and monitoring of a scalable, governed enterprise AI platform using Amazon EKS and AWS services.
- Design agentic AI workflows, autonomous agents, and multi-agent systems using LLM orchestration frameworks.
- Build secure MCP-based integrations with internal systems, vector databases, and SaaS applications such as Google Workspace and Slack.
- Implement identity, authorization, zero-trust token brokering, and least-privilege tool execution using Okta, Auth0, and JWT authorizers.
- Enforce role-based access, approval gates, and human-in-the-loop controls with deterministic policy mechanisms such as Cedar.
- Develop isolated containerized runtime environments for AI model execution, tool usage, and knowledge retrieval.
- Establish audit trails and observability for AI interactions, including cost, latency, and tool-call tracking.
- Troubleshoot AWS, Kubernetes networking, PrivateLink, network isolation, and agentic workflow issues.
- Collaborate with product, security, regulatory, privacy, quality, compliance, and business stakeholders on compliant AI infrastructure.
- Contribute to platform strategy, technology roadmaps, infrastructure-as-code practices, and long-term platform evolution.
- Mentor engineers and technical teams while promoting engineering excellence and continuous improvement.
Requirements
- Bachelor’s degree or equivalent in Computer Science, Software Engineering, Artificial Intelligence, Cloud Computing, or a related field; a Master’s degree or PhD is preferred.
- 8–12 years of relevant software development and cloud infrastructure experience with demonstrated technical leadership.
- Deep expertise in AWS architecture, Amazon EKS, Kubernetes networking, VPC, PrivateLink, IAM, KMS, and GenAI services such as AWS Bedrock.
- Experience building autonomous agents and orchestrating LLM tool-calling workflows with LangChain, LangGraph, AutoGen, or Claude Agent SDK.
- Hands-on experience with Model Context Protocol or robust governed API and tool integrations for LLMs.
- Strong experience with IAM, OAuth, JWT, and enterprise identity providers including Okta and Auth0.
- Advanced proficiency in Python, TypeScript, or Go and infrastructure-as-code tools such as Terraform or AWS CDK.
- Experience with vector databases, RAG architectures, and row-level access controls, including OpenSearch, FAISS, or pgvector.
- Experience with CI/CD pipelines, MLOps practices, Kubernetes ecosystem tools such as Helm, containerization, and observability stacks.
- Knowledge of cybersecurity and regulatory frameworks including ISO 27001, NIST, SOC 2, HIPAA, IVDD, IVDR, FDA 21 CFR 800 series, and FDA 21 CFR Part 11.
- Understanding of AI governance, software validation, data integrity, risk management, AI safety, prompt-injection defenses, and secure tool execution.
- Strong problem-solving, communication, strategic-thinking, stakeholder-influence, mentoring, adaptability, and intellectual-curiosity skills.
Benefits
- Flexible work arrangement with the ability to work from the office or home, subject to a policy requiring at least 60% or 24 hours of the work week on site.
- Role is based in Sunnyvale, California, with consideration for candidates in Durham, North Carolina; Tuesdays and Thursdays are encouraged on-site days at those campuses.
- Benefits include flexible time off or vacation, a 401(k) retirement plan with employer match, medical, dental, and vision coverage, and mindfulness programs.
- The role may be eligible for annual bonus and/or incentive compensation in addition to base pay.
