4 days ago
Pune, IndiaSenior
Responsibilities
- Design and maintain end-to-end ML pipelines for data ingestion, preprocessing, training, evaluation, deployment, and monitoring.
- Orchestrate workflows using Airflow, Prefect, Azure ML Pipelines, SageMaker Pipelines, or Vertex AI Pipelines.
- Build automated CI/CD pipelines and automate code quality, security scanning, testing, model validation, and deployment.
- Containerize and deploy AI/ML workloads using Docker, Kubernetes, serverless platforms, and cloud-native AI services.
- Implement canary, blue-green, shadow, and A/B deployment strategies for LLMs, RAG systems, and AI agents.
- Implement monitoring, observability, dashboards, alerting, incident response, and reliability practices for AI systems.
- Operate and optimize AI workloads, compute, networking, storage, security controls, and managed AI services in a major cloud platform.
- Build Infrastructure-as-Code and GenAI workflows, including embedding pipelines, vector database integrations, index refresh processes, and knowledge retrieval systems.
- Apply secure deployment, RBAC, encryption, audit logging, governance, data privacy, and Responsible AI practices.
- Optimize GPU utilization, model serving costs, token usage, storage consumption, and overall AI infrastructure expenses.
- Collaborate with data scientists, AI engineers, platform engineers, and software development teams to productionize models and integrate AI services into products.
- Create architecture diagrams, technical documentation, runbooks, SOPs, deployment guides, and on-call support documentation.
Requirements
- At least 5 years of experience in DevOps, Platform Engineering, SRE, or MLOps roles.
- At least 3 years supporting machine learning, deep learning, or AI production systems.
- Proficiency with databases, especially graph databases such as Neo4j or Memgraph, including NoSQL and SQL systems.
- Ability to perform data modeling and knowledge of embeddings, vector databases, and semantic search.
- Scripting and automation experience with Python, Golang, Bash, or similar languages.
- Strong hands-on experience with at least one major cloud platform: Microsoft Azure, AWS, or Google Cloud Platform.
- Experience deploying AI/ML workloads at scale and working with Docker, Kubernetes, and container orchestration.
- Proven experience building CI/CD pipelines and using Infrastructure-as-Code tools.
- Experience with monitoring and observability platforms.
- Working knowledge of LLMs, prompt engineering, RAG architectures, vector databases, and GenAI orchestration frameworks.
- Preferred experience with Azure OpenAI, Amazon Bedrock, Vertex AI, production LLM applications, GPU infrastructure, model evaluation frameworks, LLM observability, Responsible AI, AI governance, and security best practices.
- Relevant Azure, AWS, or GCP cloud certifications are preferred.
Benefits
- Flexible working environment.
- Volunteer time off.
- LinkedIn Learning.
- Employee Assistance Program (EAP).
Tech Stack
Apache AirflowAWSAzureBashDatabricksDatadogDockerGitHub ActionsGoGoogle Cloud PlatformGrafanaJenkinsKubernetesNeo4jPrometheusPythonTerraform
