Firmus Technologies

Principal AI Infrastructure Engineer, Kubernetes

Firmus Technologies
Apply
3 days ago
Sydney, AustraliaStaff+

Responsibilities

  • Own the Kubernetes platform reference architecture covering cluster topology, lifecycle, multi-tenancy, workload isolation, and failure domains.
  • Build backend services, APIs, controllers, operators, and automation for reliable Kubernetes cluster provisioning, configuration, upgrades, scaling, and retirement.
  • Develop repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies.
  • Design and operate cluster networking, ingress, service discovery, DNS, load balancing, network policy, service mesh, and high-performance AI networking.
  • Define persistent storage, data services, backup, restore, and disaster-recovery patterns for stateful platform and AI workloads.
  • Productionize NVIDIA GPU and network operators, device plugins, drivers, telemetry, scheduling, quotas, and topology-aware placement.
  • Establish GitOps and CI/CD practices for platform software, configuration, policy, releases, testing, rollbacks, and upgrades.
  • Build platform security through identity and access control, RBAC, secrets management, policy-as-code, image security, supply-chain controls, tenant isolation, and auditable changes.
  • Define service-level objectives and observability for metrics, logs, traces, events, capacity, and performance; diagnose complex distributed-system failures.
  • Set engineering standards and operational-readiness criteria, mentor senior engineers, resolve cross-team technical decisions, and remain hands-on in implementation.

Requirements

  • 10+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation, performance, upgrades, and control-plane failure modes.
  • Experience designing, building, and operating highly available, large-scale, multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
  • Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills and experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge covering namespaces, cgroups, systemd, kernel behavior, host networking, container runtimes, performance analysis, and low-level troubleshooting.
  • Strong Kubernetes networking expertise with Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
  • Strong infrastructure automation and GitOps experience with Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
  • Practical Kubernetes security and governance experience, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies.
  • Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, large-scale accelerator scheduling, RDMA networking, and distributed AI workloads.
  • Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
  • CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
  • Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
  • Clear technical judgment and communication, with experience influencing architecture across software, networking, security, platform, and operations teams.

Benefits

  • Full-time employment based in Australia, with locations in Sydney, NSW or Launceston, TAS.
  • Opportunity to work on sustainable AI infrastructure, GPU cloud platforms, and large-scale Kubernetes systems.

Tech Stack

AnsibleArgo CDBashElasticsearchGitHub ActionsGitLab CI/CDGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRustTerraform

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees
Contact me