3 days ago
San Mateo, CA, USAStaff+
Base Salary
$280k - $420k/yr
Responsibilities
- Lead the design and implementation of the Kubernetes-native AI Factory Reference Architecture.
- Build infrastructure for distributed training, simulation, inference, evaluation, and reinforcement learning workloads.
- Design self-service workflows that support local experimentation through large-scale distributed execution.
- Optimize shared GPU infrastructure, including scheduling, storage, networking, observability, and reliability.
- Enable dataset management, experiment tracking, artifact management, model versioning, evaluation, deployment, monitoring, and continuous model improvement.
- Develop repeatable Infrastructure as Code deployment and lifecycle-management solutions for cloud, on-premise, sovereign, and air-gapped environments.
- Evaluate AI infrastructure technologies and establish scalable, maintainable architectural patterns.
- Collaborate with ML researchers, autonomy teams, infrastructure engineers, and product teams.
Requirements
- Experience building Kubernetes-native AI or MLOps platforms for distributed machine learning workloads.
- Deep knowledge of modern AI training frameworks, including PyTorch and Hugging Face Transformers, and distributed training techniques.
- Experience operating GPU-accelerated infrastructure and distributed training systems.
- Strong understanding of Kubernetes, Linux, networking, security, storage, and distributed systems.
- Experience with GPU scheduling and large-scale AI workloads.
- Experience packaging and deploying cloud-native infrastructure with Terraform and Helm.
- Strong software engineering skills in Python, Golang, and modern cloud-native technologies.
- Experience translating ML research workflows into scalable platform capabilities.
- Preferred: experience with Ray or other distributed AI orchestration frameworks, KAI or Slurm, reinforcement learning, simulation-driven training, robotics, autonomy workloads, edge hardware, classified or air-gapped infrastructure, OpenTelemetry, Prometheus, Grafana, or open-source infrastructure projects.
Benefits
- Full-time regular employees receive pay within the listed range plus bonus, benefits, and equity.
- Temporary employees receive the listed pay range plus a temporary benefits package applicable after 60 days of employment.
- The role is a full-time regular employee position, with offers contingent on a cleared background and possible reference check.
Tech Stack
Categories
About Shield AI
Founded in 2015, Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI ’s technology actively supports operations worldwide.
