21 hours ago
Base Salary
$200k - $300k/yr
Responsibilities
- Implement and maintain AI platform configurations, orchestration tools, models, routes, guardrails, tenants, policies, role mappings, and prompt libraries.
- Design, configure, and maintain connectors and extensions to SaaS systems, data sources, and workflow tools.
- Set up and maintain logging, metrics, alerts, and other observability wiring for AI workflows and tools.
- Implement secure storage and rotation for API keys, tokens, and credentials while maintaining access-control configurations.
- Execute production model swaps, policy updates, prompt changes, version upgrades, rollout plans, change records, and rollback paths.
- Run technical evaluations and benchmarks of models, tools, and configurations and summarize performance data.
- Translate logs, metrics, and incidents into governance, risk, reliability, and compliance narratives.
- Co-author operational playbooks, runbooks, and training materials and occasionally support training or office hours.
- Serve as a technical incident point of contact by triaging issues, investigating causes, proposing mitigations, executing fixes, and documenting learnings.
Requirements
- 4–7+ years of experience in platform engineering, DevOps/SRE, ML/AI operations, or technical SaaS operations with hands-on responsibility for production systems.
- Strong fluency in APIs, integrations, infrastructure-as-config concepts, and code/JSON/YAML configuration environments with automation where appropriate.
- Hands-on experience with monitoring and observability tools, including logs, metrics, and alerts, for diagnosing issues and guiding improvements.
- Practical experience with at least one AI or automation platform class, such as LLM providers, AI productivity tools, or RPA/workflow engines.
- Experience with secure secrets management and access-control practices, including roles, permissions, key rotation, and least-privilege patterns.
- Ability to document technical work and decisions clearly for technical and non-technical audiences.
- Preferred experience includes ML/AI operations, model deployment, evaluation, drift monitoring, prompt engineering, LLM policies and guardrails, AI governance or risk frameworks, cross-functional technical risk mitigation, incident response, change management, or on-call rotations.
Benefits
- Full-time regular employee package includes pay within the listed range, bonus, benefits, and equity.
- Temporary employees receive pay within the listed range and a temporary benefits package applicable after 60 days of employment.
- Offers are contingent on a cleared background and possible reference check.
- Military fellows and part-time employees are not eligible for benefits.
Categories
About Shield AI
Shield AI builds autonomous systems for military and national security customers, combining its Hivemind autonomy software with V-BAT and X-BAT unmanned aircraft and Aechelon simulation technologies. The privately held company sells hardware, software, and related services to U.S. and allied defense agencies. Founded in 2015 and headquartered in San Diego, it operates across the U.S., Europe, the Middle East, and Asia-Pacific, and its technology is used in operational deployments.
