
ML Ops Engineer
Circadia Health6 months ago
London, United KingdomMid Level
Responsibilities
- Own and extend Apache Airflow ML pipeline orchestration for training, evaluation, deployment, retraining, validation, and promotion workflows.
- Build reproducible model deployment, versioning, rollback, shadow-release, and canary-release processes using MLflow and related tooling.
- Implement monitoring, alerting, dashboards, failure recovery, incident procedures, and runbooks for ML pipelines and production model performance.
- Deploy and operate ML workloads on AWS and support model deployment to clinical edge devices.
- Manage infrastructure as code, compute resources, storage, data transfer, cost optimization, and Snowflake integrations for ML systems.
- Establish experiment logging, artifact storage, metadata, lineage, dataset versioning, labeling provenance, and validation practices.
- Build internal tooling and templates that streamline the ML development-to-production lifecycle.
- Collaborate with ML, data, backend, firmware, embedded, and clinical teams on architecture, data quality, reliability, and technical planning.
- Ensure ML infrastructure meets healthcare security and privacy requirements, including PHI handling, audit trails, HIPAA, and SOC 2.
Requirements
- At least 4 years of experience in MLOps, ML engineering, DevOps, or a closely related infrastructure role.
- Strong Python proficiency for ML pipeline development, tooling, and automation.
- Hands-on experience with ML pipeline orchestration, particularly Apache Airflow, and with MLflow or similar model registries and experiment tracking platforms.
- Experience deploying and operating ML workloads on AWS, including Batch, EC2, S3, IAM, and CloudWatch.
- Understanding of the full ML lifecycle, including training, evaluation, deployment, monitoring, and retraining.
- Experience with Docker, infrastructure as code, Git, monitoring, logging, alerting, debugging, and incident response for distributed systems.
- Familiarity with SQL and data warehousing platforms such as Snowflake.
- Preferred experience includes edge or embedded model deployment, healthcare or medical-device systems, model serving frameworks, ML CI/CD, data versioning, distributed compute, and streaming or real-time inference.
- Preferred tools and frameworks include TorchServe, TensorFlow Serving, Triton, GitHub Actions, Jenkins, DVC, LakeFS, Apache Spark, and Dask.
- Ability to work cross-functionally, communicate clearly, take end-to-end ownership, and operate effectively in a startup environment.
Benefits
- Opportunity to work on real-world healthcare problems with measurable patient impact.
- Opportunity to build data systems powering clinical-grade AI and ML.
- Ownership in a fast-growing, mission-driven company and collaboration with a multidisciplinary team.