4 hours ago
Remote, United StatesMid Level
Base Salary
$150k - $175k/yr
Responsibilities
- Build pipelines that transform telemetry, imagery, video, sensor, and simulation data into curated and versioned training datasets.
- Develop reproducible training and evaluation workflows across cloud compute and GPU resources.
- Build and maintain infrastructure for model packaging, serving, inference, versioning, and rollback.
- Implement experiment tracking, dataset lineage, model versioning, and schema versioning across data pipelines, data lakes, and services.
- Design and operate scalable AWS infrastructure using Infrastructure as Code.
- Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving.
- Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment.
- Build evaluation frameworks, regression testing, monitoring, logging, tracing, and observability for ML systems.
- Diagnose performance, scaling, reliability, and infrastructure bottlenecks.
- Partner with Autonomy, Software, Data, Simulation, and Security teams on ML infrastructure and release processes.
- Implement secure infrastructure practices including least-privilege IAM, secrets management, access controls, and secure handling of sensitive data.
- Improve ML infrastructure reliability, security, observability, reproducibility, and cost efficiency.
Requirements
- At least 3 years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related field.
- Strong Python programming experience; Go, C++, or another systems-oriented language is preferred.
- Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure.
- Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment.
- Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code.
- Hands-on experience with Kubernetes and containerized environments.
- Strong understanding of reliability, observability, testing, automation, and maintainability.
- Ability to work across engineering disciplines and solve ambiguous technical problems with a high degree of ownership.
- U.S. citizenship and ability to obtain and maintain a U.S. Government security clearance.
- Preferred experience includes MLOps and platforms such as MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster.
- Preferred experience includes GPU or accelerator scheduling, distributed training, large-scale ML workloads, multimodal datasets, autonomy, robotics, simulation, real-time systems, or edge and embedded ML deployment.
- Experience with AWS GovCloud, GCP Assured Workloads, FedRAMP, or IL4/IL5 environments is preferred.
Benefits
- Employer-paid health, dental, and vision insurance for employees and families.
- Employer-paid life insurance.
- 401(k) program with matching.
- Unlimited PTO with an enforced two-week minimum.
- Equity package.
- Work/home office stipend.
- Global Entry benefit.
- 16 weeks of paid parental leave.
- Monthly health and wellness stipend.
Tech Stack
Categories
About HavocAI
HavocAI builds autonomous systems and software for maritime, air, and ground platforms, with a focus on uncrewed surface vessels and collaborative autonomy that lets assets share data and keep operating in contested or denied communications. The privately held company sells hardware and autonomy stacks to defense and commercial maritime customers. It was founded in 2024 and is headquartered in Providence, Rhode Island.
