11 hours ago
Base Salary
$160k - $290k/yr
Responsibilities
- Build and operate Kubernetes-native platform services, controllers, operators, deployment patterns, and runtime integrations.
- Develop reusable orchestration primitives for authoring, scheduling, and scaling pipeline work.
- Build data-processing infrastructure for storage, ingestion, validation, transformation, governance, ETL/ELT, and event-driven workflows.
- Create extensible platform capabilities covering authentication, authorization, networking, routing, secrets, and related service foundations.
- Establish reference architectures, deployment patterns, capacity guidance, benchmarks, reliability practices, and distribution approaches across cloud, on-premises, edge, and air-gapped environments.
- Advance metrics, logs, traces, dashboards, alerting, service-level objectives, diagnostics, and runbooks for platform services and workflows.
- Partner with autonomy, ML Ops, simulation, test, infrastructure, product, and customer-facing teams to turn recurring distributed-systems needs into reusable capabilities.
Requirements
- Significant experience designing and operating production distributed systems, cloud-native platforms, backend infrastructure, or data-intensive services.
- Strong software engineering skills and a record of delivering production systems in Go and Python.
- Deep understanding of distributed-systems fundamentals including failure handling, idempotency, consistency, retries, ordering, delivery semantics, backpressure, partitioning, state management, and fault tolerance.
- Experience with workflow orchestration, distributed job execution, asynchronous processing, event-driven systems, or long-running service workflows.
- Ability to define architecture and technical standards while remaining hands-on with implementation, production troubleshooting, performance analysis, and reliability improvement.
- Experience working across multiple teams to create reusable and well-documented platform capabilities.
- Clear technical communication skills for explaining complex distributed-systems architecture to specialists and downstream users.
- Preferred experience includes Kubernetes controllers and operators, Custom Resource Definitions, admission control, scheduling extensions, workload-management systems, distributed execution, workflow systems, durable messaging, event streaming, data pipelines, service networking, observability, Terraform, Helm, ArgoCD, GitOps, and infrastructure as code.
Benefits
- Full-time regular employees receive pay within the listed range, bonus, benefits, and equity.
- Temporary employees receive pay within the listed range and a temporary benefits package applicable after 60 days of employment.
- Offers are contingent on a cleared background and possible reference check.
- Military fellows and part-time employees are not eligible for benefits.
Tech Stack
About Shield AI
Shield AI builds autonomous systems for military and national security customers, combining its Hivemind autonomy software with V-BAT and X-BAT unmanned aircraft and Aechelon simulation technologies. The privately held company sells hardware, software, and related services to U.S. and allied defense agencies. Founded in 2015 and headquartered in San Diego, it operates across the U.S., Europe, the Middle East, and Asia-Pacific, and its technology is used in operational deployments.
