18 hours ago
Base Salary
$274k - $343k/yr
Responsibilities
- Lead the architecture and implementation of agentic AI systems for long-horizon reasoning, orchestration, and system-level reliability.
- Build agents for geospatial reasoning over maps and spatial data.
- Design retrieval systems for large collections of static and semi-structured documents.
- Fine-tune and evaluate embedding models for mission-critical datasets.
- Design memory systems for persistent state, long contexts, and learning from prior interactions.
- Own shared agentic infrastructure and core libraries across teams, products, and Public Sector contracts.
- Define evaluation strategies including robustness testing, failure-mode analysis, and production regression testing.
- Partner with engineering managers, product leaders, and researchers to scope initiatives and unblock execution.
- Mentor engineers and raise standards for system design, ML rigor, and production readiness.
- Travel approximately 10% for customer interactions and team needs.
Requirements
- 8+ years of experience building and deploying applied ML systems in production environments.
- Deep experience with agentic systems, autonomous workflows, or multi-step reasoning and acting ML systems.
- Strong ML systems engineering background, including model serving, pipelines, monitoring, and evaluation.
- Hands-on experience with retrieval systems, embeddings, or representation learning.
- Proficiency in Python and modern ML frameworks such as PyTorch, with end-to-end system design ability.
- Staff-level experience setting technical direction, owning ambiguous problems, and driving initiatives from 0 to 1 into production.
- Experience balancing performance, cost, reliability, and development velocity.
- Active TS security clearance.
- Preferred experience with air-gapped, classified, disconnected, on-premises, or customer data-center deployments.
- Preferred experience with DoD, intelligence community, or federal mission users.
- Preferred hands-on experience with geospatial data or GEOINT.
- Preferred experience with embedding-model training or fine-tuning, instruction tuning, LoRA/PEFT, or RLHF.
- Preferred experience building evaluation infrastructure for non-deterministic systems, including LLM-as-judge, agent regression suites, or production drift detection.
- Preferred experience turning forward-deployed prototypes into supported and documented capabilities.
Benefits
- Base salary, equity eligibility, and comprehensive health, dental, and vision coverage are provided for eligible roles.
- Retirement benefits, a learning and development stipend, and generous paid time off are offered.
- The role may be eligible for a commuter stipend.
- The position is full-time in Washington, DC, with approximately 10% travel.
- Candidates must wait 90 days before reconsideration for the same role.
Categories
About Scale AI
Scale AI builds data annotation services and AI development tools for enterprises and government agencies, sold as a platform and managed services. Its products include the Scale Generative AI Platform for building and evaluating agents and the Data Engine for collecting, curating, and labeling training data, including RLHF and model evaluation. Founded in 2016 and headquartered in San Francisco, the company is privately held and works across domains from computer vision to LLM applications.
