14 days ago
Remote, United StatesSenior
Responsibilities
- Design, build, and maintain production-grade AI systems and customer-facing AI features.
- Develop agentic workflows using LLMs, retrieval systems, tools, APIs, and backend services.
- Build backend services, orchestration systems, automation, and infrastructure for AI-powered workflows.
- Design and implement RAG systems with ingestion pipelines, embeddings, semantic retrieval, and context assembly.
- Integrate foundation models through platforms such as Amazon Bedrock and Agent Core.
- Develop prompting strategies, structured outputs, guardrails, and workflow logic for production use cases.
- Implement evaluation systems for prompts, agents, and workflows, including regression testing, trace review, golden datasets, and human QA.
- Monitor, debug, and improve production AI systems using logs, traces, evaluations, user feedback, and telemetry.
- Collaborate with engineering, product, operations, and customer-facing teams to turn ambiguous requirements into reliable systems.
- Help establish standards for testing, deployment, CI/CD, version control, code review, and operational reliability.
- Mentor and collaborate with engineers across software and AI disciplines.
- Evaluate emerging AI technologies based on business impact, maintainability, and operational reliability.
Requirements
- 5+ years of professional software engineering experience building production systems.
- Strong proficiency in Python and strong backend engineering fundamentals.
- Experience building scalable APIs, services, distributed systems, or workflow orchestration platforms.
- Hands-on experience building and shipping production AI applications using LLMs, generative AI APIs, agents, retrieval systems, or related technologies.
- Experience designing agentic workflows, tool-calling systems, structured outputs, prompt pipelines, or RAG architectures.
- Understanding of production AI challenges including hallucination mitigation, evaluation, reliability, observability, latency, and cost management.
- Experience building production software with strong standards for testing, QA, deployment, monitoring, and maintainability.
- Understanding of Git workflows, code review, CI/CD, automated testing, operational debugging, and release management.
- Experience with cloud infrastructure, preferably AWS, and with SQL and/or NoSQL databases.
- Strong debugging, systems-thinking, problem-solving, communication, and cross-functional collaboration skills.
- US citizenship or authorization to work in the US is required.
- Preferred: experience with AWS services, AI orchestration frameworks, multi-step agents, evaluation systems, vector databases, semantic retrieval, observability or LLMOps tooling, and human-in-the-loop workflows.
- Preferred: experience balancing quality, latency, reliability, and cost tradeoffs in production AI systems; mentoring engineers; and working in startup or high-ownership product environments.
- Preferred: ability to assess edge cases, failure modes, operational risk, and long-term maintainability.
