Aivar Innovations Private Limited

Senior Kubernetes Platform Engineer

Aivar Innovations Private Limited
Apply
19 days ago
Bengaluru, India or Coimbatore, IndiaSenior

Responsibilities

  • Own Kubernetes control loops, including CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines.
  • Design multi-cluster control, including workload-cluster onboarding, authentication, observation, upgrades, and outbound-only connectivity.
  • Build and operate workload-cluster agents for registration, heartbeat, inventory, command execution, reconnect behavior, versioning, rollout, credential rotation, and failure recovery.
  • Translate project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services into safe, idempotent Kubernetes reconciliation.
  • Own resource isolation and workload placement using namespaces, quotas, limits, priorities, scheduling, taints, tolerations, affinity, topology, GPU resources, and gang scheduling.
  • Integrate GPU and accelerator support, including NVIDIA GPU Operator, device plugins, MIG, node feature discovery, topology-aware scheduling, and accelerator-specific runtimes.
  • Build secure remote operations for browser-based kubectl, exec and log access, tunneled cluster connectivity, least-privilege credentials, mTLS, authorization, and audited command execution.
  • Deploy only the Kubernetes operators and services required by enabled project or cluster capabilities.
  • Design resilient systems for disconnected clusters, broken links, restarted agents, expired watches, throttled APIs, repeated reconciliation, and partially failed upgrades.
  • Work with Kubernetes API machinery, including informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.
  • Own platform upgrades, version skew, agent upgrades, CRD evolution, migrations, backward compatibility, and rollout or rollback strategies.
  • Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.
  • Review designs and code, mentor engineers, and establish patterns for reliable Kubernetes-native systems.

Requirements

  • 5+ years in software, platform, infrastructure, SRE, or cloud engineering, with substantial hands-on Kubernetes experience.
  • At least 4 years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure.
  • Strong Go skills, including production services, concurrency, interfaces, testing, profiling, and failure handling.
  • Deep understanding of Kubernetes internals, including controllers, CRDs, reconciliation, API machinery, watches and informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.
  • Experience building Kubernetes software rather than only operating clusters.
  • Experience with Kubebuilder, controller-runtime, Operator SDK, custom schedulers, admission webhooks, or Kubernetes-integrated platforms is especially valued.
  • Strong distributed-systems knowledge, including idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence.
  • Experience operating Kubernetes across cloud-managed and on-premises or private-cloud environments is valuable.
  • Networking knowledge covering TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress or gateway systems, and connectivity troubleshooting.
  • Production troubleshooting ability across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.
  • Strong security fundamentals, including service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.
  • Fluency with agentic coding tools while retaining responsibility for architecture, correctness, failure handling, and operational quality.
  • Strong pluses include multi-cluster management platforms, Kubernetes API aggregation or extension patterns, Cluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, Envoy, reverse tunnels, relay systems, secure remote cluster access, GPU scheduling, CUDA-aware workloads, distributed training infrastructure, Prometheus, OpenTelemetry, Loki, Kubernetes observability stacks, Kubernetes conformance, upgrade testing, chaos testing, and large-scale fleet management.

Tech Stack

AmbassadorArgo CDGoKubernetesPrometheusRancher

Categories

Aivar Innovations Private Limited

About Aivar Innovations Private Limited

51-200 employees
Contact me