
Sr. Engineer II, ML Ops
Alegeus Technologies LLC8 days ago
Bengaluru, IndiaSenior
Responsibilities
- Design and implement enterprise MLOps and LLMOps practices, including model and prompt versioning, CI/CD, artifact registries, environment promotion, deployment automation, retraining, rollback, and self-healing mechanisms.
- Build evaluation, monitoring, and observability frameworks for traditional ML and LLM systems, including golden datasets, automated tests, dashboards, alerts, anomaly detection, and incident response playbooks.
- Define and enforce governed AI release processes with evaluation evidence, security reviews, rollback plans, approval gates, separation of duties, audit trails, and safe rollout patterns.
- Establish playbooks, runbooks, decision frameworks, architecture reviews, and production readiness assessments while mentoring AI engineers, data scientists, and platform engineers.
- Monitor AI research and lead proof-of-concepts involving LLMs, multimodal systems, agentic workflows, advanced RAG, fine-tuning, distillation, and domain-specific model architectures.
- Design multi-agent and autonomous AI workflows using agentic frameworks and integrate advanced RAG systems with Snowflake-native data and vector databases.
- Translate applied AI research into production capabilities for claims accuracy, document processing, workflow automation, and other product opportunities.
- Shape AI platform strategy, influence product roadmaps, and mentor AI practitioners and engineering teams.
Requirements
- 7+ years of software engineering experience.
- 5+ years leading AI/ML initiatives at production scale.
- Deep hands-on expertise in at least four of MLOps, LLMOps, evaluation frameworks, drift monitoring, CI/CD for ML, and model governance.
- Strong Python skills with the ability to write, review, and debug production systems.
- Experience deploying AI in regulated environments such as healthcare, fintech, benefits administration, or equivalent.
- Proven ability to establish practices that teams adopt and operate independently.
- Preferred experience with Azure Machine Learning, MLflow, Azure DevOps, Azure OpenAI, or comparable LLM platforms.
- Experience with AutoGen, CrewAI, LangChain, Semantic Kernel, graph databases, vector databases, embedding models, RAG architectures, context engineering, or harness engineering is preferred.
- Healthcare benefits, claims processing, or consumer-directed healthcare experience is preferred.
- Experience with Docker, Kubernetes or AKS, Terraform, secrets management, and RBAC is preferred.
- Track record of converting AI research into production capabilities.
- Ability to balance operational rigor with research exploration and influence stakeholders without direct authority.
Benefits
- Flexible work environment.
- Competitive salaries, paid vacation, and holidays.
- Robust professional development programs.
- Comprehensive health, wellness, and financial packages.