4 days ago
Bengaluru, IndiaMid Level
Responsibilities
- Design, automate, and operate end-to-end machine learning training, evaluation, and deployment pipelines using Azure-native tooling.
- Manage containerized model endpoints, model versioning, traffic management, and rollback mechanisms across environments.
- Implement model performance monitoring, data drift detection, observability, and production alerting frameworks.
- Govern LLM provider relationships, API access, versioning, token consumption, and cost optimization across workloads.
- Maintain LLM gateway and prompt-versioning tooling such as LangSmith or Helicone for tracing, evaluation, and production prompt lifecycle management.
- Support reproducible experiment tracking, model cataloging, model registry processes, and governed paths from experimentation to production.
- Translate data science and AI engineering requirements into reliable, scalable production systems.
- Champion CI/CD, infrastructure as code, automated testing, and operational automation for ML workflows.
- Monitor infrastructure spending, identify optimization opportunities, and ensure platform SLAs are met.
Requirements
- 3–4 years of experience in machine learning, Azure, deployment, and pipelines.
- Experience building and operating machine learning pipelines and production model deployment infrastructure.
- Knowledge of model monitoring, data drift detection, observability, versioning, rollback, and model serving.
- Experience with LLM providers, LLM gateways, prompt versioning, tracing, evaluation, or provider governance.
- Ability to collaborate with data scientists, data engineers, and AI architects to operationalize models.
- Understanding of CI/CD, infrastructure as code, automated testing, reliability, and cost optimization for ML platforms.
