Firmus Technologies

Senior Kubernetes Platform Engineer

Firmus Technologies
Apply
1 day ago
Sydney, AustraliaSenior

Responsibilities

  • Operate and continuously improve a fleet-scale, multi-tenant Kubernetes platform across all Firmus sites.
  • Build operational tooling, guarded remediation, orchestration, and automation for estate-scale operations.
  • Execute cluster provisioning, patching, upgrades, decommissioning, version compliance, and staged or canary releases.
  • Operate and recover Kubernetes control planes, including etcd state, certificate rotation, failed upgrades, and corrupted resources.
  • Operate virtual clusters, tenant isolation patterns, GPU integration, device plugins, GPU scheduling, and driver coordination.
  • Diagnose scheduling, CNI, CSI, admission, resource contention, and other Kubernetes faults from first principles.
  • Run production Kubernetes baselines covering admission policy, workload identity, and network policy requirements.
  • Lead technical recovery during major Kubernetes incidents, drive permanent fixes, mentor engineers, and maintain runbooks and performance documentation.
  • Share the after-hours escalation roster for the Kubernetes estate.

Requirements

  • 8+ years of overall experience with substantial ownership of production Kubernetes platforms in a 24/7 environment.
  • Deep experience with Kubernetes fleet operations, cluster lifecycle, upgrades, multi-cluster management, and control-plane internals.
  • Strong production experience writing Kubernetes controllers, operators, or admission logic.
  • Experience with multi-tenant or virtual cluster patterns and Kubernetes-layer tenant isolation.
  • Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling, and driver coordination.
  • Strong infrastructure automation, infrastructure-as-code, and progressive rollout experience using tools such as OpenTofu or Terraform, Ansible, and Argo CD.
  • Strong scripting or programming skills for operational automation using Go, Python, or Bash.
  • Experience serving as a senior production escalation point, including major incident response, on-call participation, post-incident review, and runbook creation.
  • Understanding of Kubernetes security fundamentals, including admission control, workload identity, and network policy.
  • Clear technical judgment, communication, documentation, design-note, and escalation skills.
  • Preferred experience with GPU or HPC Kubernetes workloads, automated tenant or customer onboarding, vendor Kubernetes distributions, accelerated-computing reference architectures, DPU or SmartNIC networking, and open-source Kubernetes ecosystem contributions.
  • A preferred bachelor’s degree in computer science, engineering, or a related discipline, or an equivalent combination of relevant experience and training.

Benefits

  • Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
  • The role participates in a shared after-hours escalation roster within a 24/7 operations function.

Tech Stack

AnsibleArgo CDBashGoKubernetesPythonTerraform

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me