Firmus Technologies

Senior AI Infrastructure Engineer (Virtualisation)

Firmus Technologies
Apply
1 month ago
Singapore, SingaporeSenior

Responsibilities

  • Design and implement a scalable, multi-tenant control plane for AI and infrastructure workloads.
  • Develop and operate exabyte-scale S3-compatible object storage, distributed filesystems, and high-performance storage platforms.
  • Use bare-metal provisioning tools and automate hardware and cluster lifecycle management.
  • Optimize GPU, RDMA, networking, storage, and high-performance AI workload performance.
  • Develop custom Kubernetes operators and orchestration frameworks for large-scale GPU cluster commissioning.
  • Monitor, debug, validate, benchmark, and improve internal clusters and storage platforms.
  • Collaborate with SRE, site operations, and networking teams to ensure platform reliability and reproducibility.
  • Document architecture, operating procedures, performance results, and knowledge-transfer materials.
  • Participate in on-call support for production services.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 6–10 years of experience in infrastructure engineering and/or storage engineering.
  • Hands-on experience with bare-metal provisioning and software-defined storage platforms such as Ceph, Weka, Vast Data, DAOS, or Lustre.
  • Strong understanding of cloud-native infrastructure, Kubernetes, scalable systems, operating systems, computer networks, and high-performance applications.
  • Practical Linux systems engineering experience involving kernels, cgroups, system services, networking, and drivers.
  • Experience with infrastructure automation tools such as Ansible, Helm, Terraform/OpenTofu, or equivalent.
  • Understanding of firmware, BIOS, BMC/IPMI/Redfish, and low-level system tuning.
  • Proficiency in one or more of Go, Bash, Rust, or Python.
  • Experience with GPU systems, RDMA fabrics, distributed AI workloads, virtualisation, bare-metal infrastructure, and Kubernetes/Slurm platforms.
  • Strong debugging, problem-solving, documentation, and technical communication skills.
  • Experience participating in an on-call rotation supporting production services.

Benefits

  • Full-time employment.
  • Work location in Singapore or Australia, including Melbourne, Sydney, or Launceston.
  • Opportunity to work on sustainable AI infrastructure and large-scale GPU and storage platforms.

Tech Stack

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees
Contact me