3 days ago
Responsibilities
- Architect and implement the program’s MLOps strategy in alignment with project proposals and delivery roadmaps.
- Design and own enterprise-grade ML and LLM pipelines for training, validation, deployment, versioning, monitoring, and CI/CD automation.
- Build container-oriented ML platforms centered on EKS and evaluate alternative orchestration and MLOps tools.
- Implement hybrid MLOps and LLMOps workflows, including prompt-version governance, evaluation frameworks, and monitoring.
- Define standards for model deployment, monitoring, governance, automation, reliability, and scalability.
- Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML and LLM systems.
- Collaborate with data engineering, platform, DevOps, and client stakeholders on production-ready ML solutions.
- Ensure solutions meet security, governance, and compliance expectations.
- Conduct architecture reviews, troubleshoot complex ML system issues, and guide implementation across cloud-native ML platforms.
- Mentor engineers on MLOps tools, platform capabilities, and best practices.
Requirements
- 8+ years of experience in ML/AI engineering or MLOps roles with strong architecture exposure.
- Strong expertise in the AWS cloud-native ML stack, including SageMaker, EKS, Lambda, API Gateway, and CodeBuild or CodePipeline.
- Hands-on experience with at least one major MLOps toolset and awareness of alternatives including MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, and Seldon.
- Deep understanding of model lifecycle management, including feature engineering, training, model registries, deployment, monitoring, and governance.
- Experience implementing LLMOps pipelines with prompt versioning, evaluation metrics, and automation frameworks.
- Experience implementing ML CI/CD pipelines for automated training, testing, validation, model promotion, and endpoint deployment.
- Experience with infrastructure-as-code tools, CI/CD pipelines, Kubernetes-based development, feature engineering pipelines, and Feature Store management.
- Understanding of lineage tracking, reproducibility, data snapshots, feature versions, code versioning, and metadata tracking.
- Hands-on experience with AWS Bedrock and Agentcore.
- Experience with CloudWatch, SageMaker Model Monitor, Prometheus, and Grafana.
- Strong Python and cloud-native development skills.
- Understanding of security best practices, IAM, secrets management, and artifact governance.
- Preferred experience with vector databases, RAG pipelines, multi-agent AI systems, Terraform, Helm, CDK, model drift detection, A/B testing, canary rollouts, blue-green deployments, and OpenTelemetry.
- Preferred SQL and data transformation experience using Snowflake, Databricks, or Spark.
- Ability to translate business goals into scalable AI/ML platform designs and guide engineering teams through technical uncertainty and design choices.
- Strong communication and cross-team collaboration skills.
Tech Stack
Apache AirflowApache SparkAWSDatabricksGrafanaHelmKubernetesMLflowPrometheusPythonSnowflakeSQLTerraform
Categories
About Quantiphi
Quantiphi is an AI-first digital engineering and consulting firm that designs and implements machine learning, data, and cloud solutions for large enterprises across industries. Its teams build production systems—such as generative AI applications, computer vision, and predictive analytics—primarily on Google Cloud and other hyperscalers, delivered as professional services and managed solutions. Founded in 2013 and headquartered in Marlborough, Massachusetts, the company is privately held and a Google Cloud partner.
