Firmus Technologies

Senior AI Infrastructure Engineer, Kubernetes

Firmus Technologies
Apply
4 hours ago

Responsibilities

  • Own the Kubernetes platform reference architecture for management and workload clusters, including lifecycle, multi-tenancy, workload isolation, and failure-domain design.
  • Build backend services, APIs, controllers, operators, and automation for provisioning, configuring, upgrading, scaling, and retiring Kubernetes clusters.
  • Develop repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies.
  • Design and operate cluster networking, including CNI, ingress, service discovery, DNS, load balancing, network policy, service mesh, and high-performance networking integrations.
  • Define persistent-storage, data-service, backup, restore, and disaster-recovery patterns for stateful platform and AI workloads.
  • Integrate and productionize NVIDIA GPU and Network Operators, device plugins, drivers, telemetry, scheduling, quotas, and topology-aware placement.
  • Establish GitOps and CI/CD practices for platform software, configuration, policy, and releases, including testing, rollout, rollback, and upgrade procedures.
  • Build security into the platform through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable changes.
  • Define service-level objectives and observability capabilities, diagnose complex distributed-systems failures, and reduce recurring operational toil.
  • Set engineering standards and readiness criteria, mentor senior engineers, resolve cross-team technical decisions, and remain directly involved in implementation.

Requirements

  • 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms.
  • At least 3 years operating at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, cluster performance, upgrades, reconciliation patterns, and control-plane failure modes.
  • Experience designing, building, and operating highly available, large-scale, multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
  • Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills and experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel behavior, host networking, container runtimes, performance analysis, and low-level troubleshooting.
  • Strong Kubernetes networking expertise with Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
  • Strong infrastructure automation and GitOps experience with Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
  • Experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies.
  • Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, large-scale accelerator scheduling, RDMA networking, and distributed AI workloads.
  • Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
  • CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
  • Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
  • Clear technical judgment and communication, with experience influencing architecture across software, networking, security, platform, and operations teams.

Benefits

  • Full-time employment.
  • Role is based in the San Francisco Bay Area.
  • Reports to the Head of AI Platform.
  • Opportunity to work on sustainable, large-scale AI infrastructure and GPU cloud technology.
  • Inclusive workplace committed to diversity and equal opportunity.

Tech Stack

AnsibleArgo CDBashDockerElasticsearchGitHub ActionsGitLab CI/CDGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRustTerraform

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me