4 hours ago
Remote, United StatesStaff+
Base Salary
$175k - $200k/yr
Responsibilities
- Design and iterate on agent behavior across live, long-horizon, multi-turn, and multi-agent workflows.
- Design retrieval, context, memory, and state architectures that keep agents grounded in source data.
- Create prompt and context templates using few-shot examples, structured formatting, and reasoning scaffolding.
- Improve agent performance through prompting, tool-use strategy, and context construction validated by experimentation.
- Build production-grounded evaluations to measure performance, regressions, failure modes, and edge cases.
- Author severity- and cost-weighted evaluation rubrics, quality heuristics, and thresholds.
- Design and validate confidence- and uncertainty-based escalation paths to human review.
- Optimize cost-aware performance across latency, reliability, accuracy, context construction, and tool-call economy.
- Evaluate model changes, run comparative baselines, and make go/no-go launch decisions.
- Maintain product-level AI documentation including model cards, intended use, limitations, and failure modes.
- Partner with Product, product managers, and Engineering to make agents steerable, trustworthy, and ready to scale.
Requirements
- Equivalent practical experience demonstrating the depth required for the staff-level role is valued.
- 8+ years of production software engineering experience are required, including 3+ years of hands-on ownership of ML, LLM, or agentic systems in production.
- Production experience with ML, LLM, or agentic systems must include direct experience in healthcare, finance, or another regulated industry.
- Ability to diagnose agent failures and attribute fixes to instruction, retrieval, context, or memory design.
- Ability to weigh failures by severity and cost rather than frequency alone.
- Hands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic.
- Working familiarity with AWS AI/ML services, including Bedrock and SageMaker.
- Evidence-led judgment, credibility to challenge launch decisions, and a hands-on experimentation and building approach.
- Preferred: experience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty affect care teams or patients.
- Preferred: experience with long-horizon, multi-turn, or multi-agent workflows and product-level AI documentation such as model cards.
Benefits
- Flexible, remote-friendly work culture.
- Employee-driven programs and initiatives supporting personal and professional development.
- Membership in a diverse, purpose-driven Arcadian community.
- Opportunity to define agent performance, safety, and readiness standards for production healthcare workflows.
- Meaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scale.
Tech Stack
Categories
About Arcadia
Arcadia is the global utility data and energy solutions platform. With our leading data platform, AI-powered analytics, industry expertise, and expansive partner network, we deliver solutions for every stage of the enterprise energy management lifecycle across carbon, cost, and reliability. Join us at www.arcadia.com/careers.