Senior Kubernetes Platform Engineer
Firmus Technologies1 day ago
Sydney, AustraliaSenior
Responsibilities
- Operate and continuously improve a fleet-scale, multi-tenant Kubernetes platform across all Firmus sites.
- Build operational tooling, guarded remediation, orchestration, and automation for estate-scale operations.
- Execute cluster provisioning, patching, upgrades, decommissioning, version compliance, and staged or canary releases.
- Operate and recover Kubernetes control planes, including etcd state, certificate rotation, failed upgrades, and corrupted resources.
- Operate virtual clusters, tenant isolation patterns, GPU integration, device plugins, GPU scheduling, and driver coordination.
- Diagnose scheduling, CNI, CSI, admission, resource contention, and other Kubernetes faults from first principles.
- Run production Kubernetes baselines covering admission policy, workload identity, and network policy requirements.
- Lead technical recovery during major Kubernetes incidents, drive permanent fixes, mentor engineers, and maintain runbooks and performance documentation.
- Share the after-hours escalation roster for the Kubernetes estate.
Requirements
- 8+ years of overall experience with substantial ownership of production Kubernetes platforms in a 24/7 environment.
- Deep experience with Kubernetes fleet operations, cluster lifecycle, upgrades, multi-cluster management, and control-plane internals.
- Strong production experience writing Kubernetes controllers, operators, or admission logic.
- Experience with multi-tenant or virtual cluster patterns and Kubernetes-layer tenant isolation.
- Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling, and driver coordination.
- Strong infrastructure automation, infrastructure-as-code, and progressive rollout experience using tools such as OpenTofu or Terraform, Ansible, and Argo CD.
- Strong scripting or programming skills for operational automation using Go, Python, or Bash.
- Experience serving as a senior production escalation point, including major incident response, on-call participation, post-incident review, and runbook creation.
- Understanding of Kubernetes security fundamentals, including admission control, workload identity, and network policy.
- Clear technical judgment, communication, documentation, design-note, and escalation skills.
- Preferred experience with GPU or HPC Kubernetes workloads, automated tenant or customer onboarding, vendor Kubernetes distributions, accelerated-computing reference architectures, DPU or SmartNIC networking, and open-source Kubernetes ecosystem contributions.
- A preferred bachelor’s degree in computer science, engineering, or a related discipline, or an equivalent combination of relevant experience and training.
Benefits
- Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
- The role participates in a shared after-hours escalation roster within a 24/7 operations function.
Tech Stack
Categories
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.