1 month ago
Pune, India +2 moreStaff+

Responsibilities

  • Architect and build scalable LLMOps platforms for enterprise-grade GenAI systems.
  • Design and manage end-to-end LLM pipelines covering data ingestion, embeddings, evaluation, fine-tuning, and inference.
  • Develop, deploy, monitor, and optimize ML and LLM systems for production at scale.
  • Build agentic AI workflows and operational capabilities including orchestration, evaluation, observability, and reflection loops.
  • Implement guardrails, output filtering, context-aware routing, evaluation harnesses, metrics logging, and incident response.
  • Create scalable Kubernetes- and GPU-aware deployment frameworks for LLM inference across cloud and hybrid environments.
  • Drive platform automation using Docker, Kubernetes, GitOps, Terraform, and DevOps practices.
  • Design systems for prompt and model versioning, AI governance, automated testing, and prompt quality scoring.
  • Lead ideation, prototyping, and scaling of reusable internal LLMOps accelerators.
  • Provide technical leadership and lead development teams.

Requirements

  • 10–14 years of experience working on ML projects, including product building, hands-on engineering, technical leadership, and leading development teams.
  • Experience with model development, training, deployment at scale, and production performance monitoring.
  • Strong Python, data engineering, FastAPI, and NLP skills.
  • Knowledge of LangChain, LlamaIndex, Langtrace, Langfuse, LLM evaluation, MLflow, and BentoML.
  • Experience with proprietary and open-source LLMs and LLM fine-tuning, including PEFT and CPT.
  • Experience creating agentic AI workflows using CrewAI, LangGraph, AutoGen, and Semantic Kernel.
  • Experience with performance optimization, RAG, guardrails, AI governance, prompt engineering, evaluation, and observability.
  • Experience deploying GenAI applications at scale in cloud and on-premises environments using DevOps and MLOps practices.
  • Working knowledge of Kubernetes and Terraform and experience with at least one of AWS, GCP, or Azure.
  • Experience with Ray, Truss, OpenAI Evals, Ragas, Rebuff, Outlines, Helm, Docker, Airflow, Prefect, Feast, Spark, Flink, Parquet, Delta, Azure ML, Vertex AI, AWS Bedrock, or SageMaker.
  • Experience with LLMOps, prompt evaluation, cloud-native deployment, ML pipelines, data platforms, and scalable inference architectures.
  • Python is required; Bash, YAML, and Terraform HCL are preferred.
  • Strong communication, presentation, collaboration, and product-thinking skills.
Fractal Analytics

About Fractal Analytics

5,001-10,000 employees
Contact me