Aivar Innovations Private Limited

Senior MLOps / AI Platform Engineer

Aivar Innovations Private Limited
Apply
1 month ago
Coimbatore, IndiaSenior / Staff+

Responsibilities

  • Design and build an enterprise-grade MLOps and AIOps platform running on Kubernetes.
  • Develop infrastructure and platform capabilities to deploy, manage, scale, and observe machine learning, deep learning, and generative AI workloads.
  • Write Kubernetes Operators and Custom Resource Definitions.
  • Design and operate cloud-native infrastructure on AWS, including GPU-accelerated workloads.
  • Deploy, operate, and troubleshoot production machine learning, deep learning, and generative AI models.
  • Build production-grade APIs, controllers, and distributed backend services using Go or Python.
  • Operate model-serving frameworks and optimize distributed inference, batching, quantization, and performance.
  • Lead major technical initiatives from architecture through production deployment and help shape platform architecture and roadmap.
  • Collaborate with Product, Engineering, AI/ML, DevOps, Customer Delivery, and Leadership teams.

Requirements

  • 5–8 years of experience in software engineering, platform engineering, DevOps, SRE, MLOps, or related infrastructure roles.
  • Strong hands-on Kubernetes experience, including Operators and Custom Resource Definitions using Kubebuilder, Operator SDK, or equivalent.
  • Experience designing and operating AWS infrastructure, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch.
  • Experience deploying and operating machine learning, deep learning, or generative AI models in production.
  • Experience running and troubleshooting GPU-accelerated Kubernetes workloads, including GPU scheduling, utilization, memory constraints, and performance.
  • Familiarity with KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent model-serving technologies.
  • Strong programming experience in Go or Python and experience building production-grade APIs, controllers, or distributed backend services.
  • Experience with containers, Helm, CI/CD, infrastructure as code, Prometheus, OpenTelemetry, and Grafana.
  • Strong understanding of Linux, networking, storage, security, and distributed-system fundamentals.
  • Preferred experience building MLOps platforms, AI infrastructure platforms, internal developer platforms, or Kubernetes-based enterprise products.
  • Preferred experience with CUDA C/C++ or Triton for GPU kernels or performance-critical code.
  • Preferred experience with large language model serving, distributed inference, batching, quantization, or inference-performance optimization.
  • Preferred experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation, AWS Inferentia, Trainium, SageMaker, or Amazon Bedrock.
  • Preferred experience operating AI platforms across hybrid-cloud, on-premises, airgapped, or multi-tenant environments.

Benefits

  • Work on a core Kubernetes-native AI platform enabling enterprises to operate AI workloads at scale.
  • Solve challenging infrastructure problems involving Kubernetes, AWS, GPUs, distributed systems, model serving, and enterprise AI operations.
  • Influence product direction and platform architecture in collaboration with Product, Engineering, and Leadership.
  • Take ownership of major technical initiatives with visible impact at a fast-growing AI startup.
  • Build a career with accelerated growth and lasting influence from technical decisions.

Tech Stack

AWSGoGrafanaHelmKubernetesLinuxPrometheusPython
Aivar Innovations Private Limited

About Aivar Innovations Private Limited

51-200 employees
Contact me