
Lead Architect
Fractal Analytics1 month ago
Pune, India +2 moreStaff+
Responsibilities
- Architect and build scalable LLMOps platforms for enterprise-grade GenAI systems.
- Design and manage end-to-end LLM pipelines covering data ingestion, embeddings, evaluation, fine-tuning, and inference.
- Develop, deploy, monitor, and optimize ML and LLM systems for production at scale.
- Build agentic AI workflows and operational capabilities including orchestration, evaluation, observability, and reflection loops.
- Implement guardrails, output filtering, context-aware routing, evaluation harnesses, metrics logging, and incident response.
- Create scalable Kubernetes- and GPU-aware deployment frameworks for LLM inference across cloud and hybrid environments.
- Drive platform automation using Docker, Kubernetes, GitOps, Terraform, and DevOps practices.
- Design systems for prompt and model versioning, AI governance, automated testing, and prompt quality scoring.
- Lead ideation, prototyping, and scaling of reusable internal LLMOps accelerators.
- Provide technical leadership and lead development teams.
Requirements
- 10–14 years of experience working on ML projects, including product building, hands-on engineering, technical leadership, and leading development teams.
- Experience with model development, training, deployment at scale, and production performance monitoring.
- Strong Python, data engineering, FastAPI, and NLP skills.
- Knowledge of LangChain, LlamaIndex, Langtrace, Langfuse, LLM evaluation, MLflow, and BentoML.
- Experience with proprietary and open-source LLMs and LLM fine-tuning, including PEFT and CPT.
- Experience creating agentic AI workflows using CrewAI, LangGraph, AutoGen, and Semantic Kernel.
- Experience with performance optimization, RAG, guardrails, AI governance, prompt engineering, evaluation, and observability.
- Experience deploying GenAI applications at scale in cloud and on-premises environments using DevOps and MLOps practices.
- Working knowledge of Kubernetes and Terraform and experience with at least one of AWS, GCP, or Azure.
- Experience with Ray, Truss, OpenAI Evals, Ragas, Rebuff, Outlines, Helm, Docker, Airflow, Prefect, Feast, Spark, Flink, Parquet, Delta, Azure ML, Vertex AI, AWS Bedrock, or SageMaker.
- Experience with LLMOps, prompt evaluation, cloud-native deployment, ML pipelines, data platforms, and scalable inference architectures.
- Python is required; Bash, YAML, and Terraform HCL are preferred.
- Strong communication, presentation, collaboration, and product-thinking skills.
Tech Stack
Apache AirflowApache FlinkApache SparkAWSAzureBashDockerFastAPIGoogle Cloud PlatformHelmKubernetesMLflowPythonTerraform