7 hours ago
Remote, United StatesMid Level
Base Salary
$150k - $185k/yr
Responsibilities
- Build and maintain infrastructure and scalable pipelines for video, imagery, telemetry, sensor, autonomy-log, mission, and field-test data.
- Own data-lake ingestion, storage, indexing, metadata, access patterns, and lifecycle management.
- Create curated, searchable, versioned, and reproducible datasets for ML training, evaluation, debugging, benchmarking, and regression testing.
- Develop tools and workflows for data search, filtering, tagging, retrieval, annotation, synchronization, replay, visualization, and analysis.
- Build automated data-quality checks, lineage and reproducibility standards, monitoring, and observability for critical pipelines.
- Partner with Autonomy, Perception, Software, Simulation, Field Operations, and Program teams to translate engineering and ML requirements into data capabilities.
- Troubleshoot complex data and infrastructure issues and drive them through resolution.
Requirements
- Bachelor’s degree in Computer Science, Data Science, Machine Learning, Electrical Engineering, Computer Engineering, Robotics, Applied Mathematics, or a related technical field.
- At least three years of experience in data engineering, ML infrastructure, data platforms, backend systems, MLOps, or related engineering roles.
- Experience designing and operating production data pipelines for large-scale structured, semi-structured, or unstructured datasets.
- Experience with video, imagery, time-series telemetry, sensor data, logs, or other high-volume operational data.
- Strong programming skills in Python and SQL.
- Experience with cloud storage, object stores, data lakes, databases, distributed processing, or modern data platforms.
- Familiarity with dataset versioning, metadata management, data lineage, access controls, reproducible data workflows, testing, reliability, maintainability, and observability.
- Strong debugging skills and the ability to work independently across complex data pipelines and production infrastructure.
- U.S. citizenship and the ability to obtain and maintain a U.S. Government security clearance.
- Preferred experience includes ML infrastructure, MLOps, training pipelines, experiment tracking, model evaluation, model registries, autonomous-system datasets, search and replay tools, annotation workflows, sensor synchronization, and multimodal dataset construction.
- Preferred experience also includes S3-compatible storage, PostgreSQL, Spark, Ray, Airflow, Dagster, Kubernetes, Docker, Kafka, data catalogs, dataset-versioning platforms, feature stores, labeling tools, and defense, robotics, autonomy, aerospace, or dual-use technology programs.
Benefits
- Employer-paid health, dental, and vision insurance for employees and families.
- Employer-paid life insurance.
- 401(k) program with matching.
- Unlimited PTO with an enforced two-week minimum.
- Equity package.
- Work/home office stipend.
- Global Entry.
- 16 weeks of paid parental leave.
- Monthly health and wellness stipend.
Tech Stack
Categories
Data EngineeringML Engineering
About HavocAI
HavocAI builds autonomous systems and software for maritime, air, and ground platforms, with a focus on uncrewed surface vessels and collaborative autonomy that lets assets share data and keep operating in contested or denied communications. The privately held company sells hardware and autonomy stacks to defense and commercial maritime customers. It was founded in 2024 and is headquartered in Providence, Rhode Island.
