Principal AI Infrastructure Engineer, Kubernetes
Firmus Technologies3 days ago
Sydney, AustraliaStaff+
Responsibilities
- Own the Kubernetes platform reference architecture covering cluster topology, lifecycle, multi-tenancy, workload isolation, and failure domains.
- Build backend services, APIs, controllers, operators, and automation for reliable Kubernetes cluster provisioning, configuration, upgrades, scaling, and retirement.
- Develop repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies.
- Design and operate cluster networking, ingress, service discovery, DNS, load balancing, network policy, service mesh, and high-performance AI networking.
- Define persistent storage, data services, backup, restore, and disaster-recovery patterns for stateful platform and AI workloads.
- Productionize NVIDIA GPU and network operators, device plugins, drivers, telemetry, scheduling, quotas, and topology-aware placement.
- Establish GitOps and CI/CD practices for platform software, configuration, policy, releases, testing, rollbacks, and upgrades.
- Build platform security through identity and access control, RBAC, secrets management, policy-as-code, image security, supply-chain controls, tenant isolation, and auditable changes.
- Define service-level objectives and observability for metrics, logs, traces, events, capacity, and performance; diagnose complex distributed-system failures.
- Set engineering standards and operational-readiness criteria, mentor senior engineers, resolve cross-team technical decisions, and remain hands-on in implementation.
Requirements
- 10+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years at senior staff, principal, or equivalent level.
- Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation, performance, upgrades, and control-plane failure modes.
- Experience designing, building, and operating highly available, large-scale, multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
- Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills and experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
- Expert Linux systems knowledge covering namespaces, cgroups, systemd, kernel behavior, host networking, container runtimes, performance analysis, and low-level troubleshooting.
- Strong Kubernetes networking expertise with Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
- Strong infrastructure automation and GitOps experience with Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
- Practical Kubernetes security and governance experience, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
- Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies.
- Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, large-scale accelerator scheduling, RDMA networking, and distributed AI workloads.
- Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
- CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
- Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
- Clear technical judgment and communication, with experience influencing architecture across software, networking, security, platform, and operations teams.
Benefits
- Full-time employment based in Australia, with locations in Sydney, NSW or Launceston, TAS.
- Opportunity to work on sustainable AI infrastructure, GPU cloud platforms, and large-scale Kubernetes systems.