10 days ago
Responsibilities
- Own the technical direction, architecture, interfaces, and engineering standards for the MLOps platform.
- Design and operate dataset versioning, lineage, splits, reproducibility, and labeling-integration workflows.
- Build and operate a model registry containing versioned artifacts, metadata, evaluation results, lineage, and approval workflows.
- Define offline benchmarks, simulation rollouts, policy-gating harnesses, and metrics used to qualify models.
- Own the path from registered models to inference on Apollo, including packaging, on-robot versioning, rollback, and observability.
- Coordinate the deploy-and-telemetry connection with Connect and Data Platform.
- Lead technical projects from architecture through implementation, validation, and iteration.
- Mentor MLOps engineers through code reviews, design reviews, and direct collaboration.
- Establish platform contracts and standards across MLOps, Autonomy, Data Platform, TeleOp, and Training Infrastructure.
Requirements
- Deep proficiency in Python and at least one systems-level language: Go, Rust, or C++.
- 8+ years of professional software engineering experience in ML platforms or related infrastructure, or 4+ years of hands-on experience owning an MLOps platform that shipped models to production.
- Proven experience owning and delivering an end-to-end production MLOps platform covering datasets, experiment tracking, model registry, evaluation, and serving.
- Experience with dataset versioning using DVC, LakeFS, Delta, or an equivalent; experiment tracking using MLflow, W&B, or Determined; model registries; and policy serving.
- Strong experience designing service-oriented systems on Kubernetes and working across platform APIs and compute infrastructure.
- Experience defining ML model evaluation and qualification frameworks where regressions have high costs, such as robotics, safety-critical, or customer-facing production systems.
- Demonstrated ability to lead technical projects, influence cross-functional teams, establish adopted standards, and mentor engineers.
- Proficiency with AWS, GCP, or Azure, Docker, Git, and modern CI/CD workflows.
- Preferred experience deploying models to edge or embedded targets, including on-device inference, ONNX Runtime, TensorRT, or robot fleets.
- Preferred experience with reinforcement learning training and evaluation infrastructure, embodied agents, rollout workers, replay buffers, and simulation-evaluation harnesses.
- Familiarity with humanoid robotics, dexterous manipulation, teleoperation data, policy gating, shadow deployments, or staged autonomy rollouts.
- Experience with IsaacSim, MuJoCo, or equivalent simulation-in-the-loop evaluation.
- Open-source contributions to MLOps tooling such as MLflow, BentoML, KServe, or Ray Serve.
- A bachelor’s degree in Computer Science, Machine Learning, or a related technical field is considered; a master’s degree is preferred.
Benefits
- Direct-hire position.
- Equal employment opportunity and anti-discrimination protections are provided.
- The role requires prolonged desk and computer work, occasional lifting of up to 15 pounds, and the ability to read screens and communicate by hearing and speech.
Tech Stack
Categories
About Apptronik
Apptronik is a human-centered robotics company developing AI-powered robots to support humanity in every facet of life. Our humanoid robot, Apollo, is designed to collaborate thoughtfully with humans—initially in critical industries such as manufacturing and logistics, with future applications in healthcare, the home, and beyond. Apollo is the culmination of nearly a decade of development, drawing on Apptronik’s extensive work on 15 previous robots, including NASA’s Valkyrie robot. Apptronik started out of the Human Centered Robotics Lab at the University of Texas at Austin and has more than 350 employees. Learn more at apptronik.com
