Senior AI Infrastructure Engineer, Kubernetes
Firmus Technologies4 hours ago
Responsibilities
- Own the Kubernetes platform reference architecture for management and workload clusters, including lifecycle, multi-tenancy, workload isolation, and failure-domain design.
- Build backend services, APIs, controllers, operators, and automation for provisioning, configuring, upgrading, scaling, and retiring Kubernetes clusters.
- Develop repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies.
- Design and operate cluster networking, including CNI, ingress, service discovery, DNS, load balancing, network policy, service mesh, and high-performance networking integrations.
- Define persistent-storage, data-service, backup, restore, and disaster-recovery patterns for stateful platform and AI workloads.
- Integrate and productionize NVIDIA GPU and Network Operators, device plugins, drivers, telemetry, scheduling, quotas, and topology-aware placement.
- Establish GitOps and CI/CD practices for platform software, configuration, policy, and releases, including testing, rollout, rollback, and upgrade procedures.
- Build security into the platform through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable changes.
- Define service-level objectives and observability capabilities, diagnose complex distributed-systems failures, and reduce recurring operational toil.
- Set engineering standards and readiness criteria, mentor senior engineers, resolve cross-team technical decisions, and remain directly involved in implementation.
Requirements
- 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms.
- At least 3 years operating at senior staff, principal, or equivalent level.
- Deep knowledge of Kubernetes internals, cluster performance, upgrades, reconciliation patterns, and control-plane failure modes.
- Experience designing, building, and operating highly available, large-scale, multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
- Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills and experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
- Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel behavior, host networking, container runtimes, performance analysis, and low-level troubleshooting.
- Strong Kubernetes networking expertise with Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
- Strong infrastructure automation and GitOps experience with Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
- Experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
- Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies.
- Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, large-scale accelerator scheduling, RDMA networking, and distributed AI workloads.
- Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
- CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
- Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
- Clear technical judgment and communication, with experience influencing architecture across software, networking, security, platform, and operations teams.
Benefits
- Full-time employment.
- Role is based in the San Francisco Bay Area.
- Reports to the Head of AI Platform.
- Opportunity to work on sustainable, large-scale AI infrastructure and GPU cloud technology.
- Inclusive workplace committed to diversity and equal opportunity.
Tech Stack
Categories
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.