Senior MLOps / AI Platform Engineer
Aivar Innovations Private Limited1 month ago
Coimbatore, IndiaSenior / Staff+
Responsibilities
- Design and build an enterprise-grade MLOps and AIOps platform running on Kubernetes.
- Develop infrastructure and platform capabilities to deploy, manage, scale, and observe machine learning, deep learning, and generative AI workloads.
- Write Kubernetes Operators and Custom Resource Definitions.
- Design and operate cloud-native infrastructure on AWS, including GPU-accelerated workloads.
- Deploy, operate, and troubleshoot production machine learning, deep learning, and generative AI models.
- Build production-grade APIs, controllers, and distributed backend services using Go or Python.
- Operate model-serving frameworks and optimize distributed inference, batching, quantization, and performance.
- Lead major technical initiatives from architecture through production deployment and help shape platform architecture and roadmap.
- Collaborate with Product, Engineering, AI/ML, DevOps, Customer Delivery, and Leadership teams.
Requirements
- 5–8 years of experience in software engineering, platform engineering, DevOps, SRE, MLOps, or related infrastructure roles.
- Strong hands-on Kubernetes experience, including Operators and Custom Resource Definitions using Kubebuilder, Operator SDK, or equivalent.
- Experience designing and operating AWS infrastructure, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch.
- Experience deploying and operating machine learning, deep learning, or generative AI models in production.
- Experience running and troubleshooting GPU-accelerated Kubernetes workloads, including GPU scheduling, utilization, memory constraints, and performance.
- Familiarity with KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent model-serving technologies.
- Strong programming experience in Go or Python and experience building production-grade APIs, controllers, or distributed backend services.
- Experience with containers, Helm, CI/CD, infrastructure as code, Prometheus, OpenTelemetry, and Grafana.
- Strong understanding of Linux, networking, storage, security, and distributed-system fundamentals.
- Preferred experience building MLOps platforms, AI infrastructure platforms, internal developer platforms, or Kubernetes-based enterprise products.
- Preferred experience with CUDA C/C++ or Triton for GPU kernels or performance-critical code.
- Preferred experience with large language model serving, distributed inference, batching, quantization, or inference-performance optimization.
- Preferred experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation, AWS Inferentia, Trainium, SageMaker, or Amazon Bedrock.
- Preferred experience operating AI platforms across hybrid-cloud, on-premises, airgapped, or multi-tenant environments.
Benefits
- Work on a core Kubernetes-native AI platform enabling enterprises to operate AI workloads at scale.
- Solve challenging infrastructure problems involving Kubernetes, AWS, GPUs, distributed systems, model serving, and enterprise AI operations.
- Influence product direction and platform architecture in collaboration with Product, Engineering, and Leadership.
- Take ownership of major technical initiatives with visible impact at a fast-growing AI startup.
- Build a career with accelerated growth and lasting influence from technical decisions.