Senior Kubernetes Platform Engineer
Aivar Innovations Private Limited19 days ago
Bengaluru, India or Coimbatore, IndiaSenior
Responsibilities
- Own Kubernetes control loops by designing and building CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines using Go and controller-runtime.
- Design and implement multi-cluster control, including workload-cluster onboarding, authentication, observation, upgrades, and control over outbound-only connections.
- Build the workload-cluster agent for registration, heartbeats, inventory, command execution, reconnect behavior, versioning, rollouts, credential rotation, and failure recovery.
- Translate platform intent into safe, idempotent Kubernetes reconciliation for project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services.
- Own resource isolation and workload placement using namespaces, quotas, limits, priority, scheduling, taints, tolerations, affinity, topology, GPU resources, and gang scheduling.
- Integrate GPU and accelerator support, including NVIDIA GPU Operator, device plugins, MIG, node feature discovery, topology-aware scheduling, and accelerator-specific runtimes.
- Build secure remote operations with browser-based kubectl, exec and log access, tunneled connectivity, least-privilege credentials, mTLS, authorization, and auditable command execution.
- Build capability deployment mechanisms that install only the required operators and services on enabled projects or clusters.
- Design resilient systems for disconnected clusters, broken links, restarted agents, expired watches, throttled APIs, repeated reconciliation, and partially failed upgrades.
- Work with Kubernetes API machinery, including informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.
- Own safe platform upgrades, version skew, agent upgrades, CRD evolution, migration, backward compatibility, and rollout and rollback strategies.
- Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.
- Review designs and code, mentor engineers, and establish patterns for reliable Kubernetes-native systems.
Requirements
- 5+ years of experience in software, platform, infrastructure, SRE, or cloud engineering, including at least 4 years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure.
- Strong Go experience designing production services with concurrency, interfaces, testing, profiling, and failure handling.
- Deep knowledge of Kubernetes internals, including controllers, CRDs, reconciliation, API machinery, watches and informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.
- Demonstrated experience building Kubernetes software rather than only operating clusters.
- Experience with or interest in Kubebuilder, controller-runtime, Operator SDK, custom schedulers, admission webhooks, or Kubernetes-integrated platforms.
- Strong distributed-systems understanding, including idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence.
- Experience operating Kubernetes across cloud-managed and on-premises or private-cloud environments.
- Networking knowledge covering TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress or gateways, and connectivity troubleshooting.
- Production troubleshooting ability across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.
- Strong security fundamentals covering service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.
- Fluency with agentic coding tools while retaining responsibility for architecture, correctness, failure handling, and operational quality.
- Preferred experience with multi-cluster management platforms, Kubernetes API aggregation or extension patterns, Cluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, Envoy, reverse tunnels, relay systems, GPU scheduling, NVIDIA GPU Operator, MIG, CUDA-aware workloads, distributed training infrastructure, Prometheus, OpenTelemetry, Loki, Kubernetes observability stacks, conformance testing, upgrade testing, chaos testing, or large-scale fleet management.