Senior AI Infrastructure Engineer (Virtualisation)
Firmus Technologies1 month ago
Singapore, SingaporeSenior
Responsibilities
- Design and implement a scalable, multi-tenant control plane for AI and infrastructure workloads.
- Develop and operate exabyte-scale S3-compatible object storage, distributed filesystems, and high-performance storage platforms.
- Use bare-metal provisioning tools and automate hardware and cluster lifecycle management.
- Optimize GPU, RDMA, networking, storage, and high-performance AI workload performance.
- Develop custom Kubernetes operators and orchestration frameworks for large-scale GPU cluster commissioning.
- Monitor, debug, validate, benchmark, and improve internal clusters and storage platforms.
- Collaborate with SRE, site operations, and networking teams to ensure platform reliability and reproducibility.
- Document architecture, operating procedures, performance results, and knowledge-transfer materials.
- Participate in on-call support for production services.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 6–10 years of experience in infrastructure engineering and/or storage engineering.
- Hands-on experience with bare-metal provisioning and software-defined storage platforms such as Ceph, Weka, Vast Data, DAOS, or Lustre.
- Strong understanding of cloud-native infrastructure, Kubernetes, scalable systems, operating systems, computer networks, and high-performance applications.
- Practical Linux systems engineering experience involving kernels, cgroups, system services, networking, and drivers.
- Experience with infrastructure automation tools such as Ansible, Helm, Terraform/OpenTofu, or equivalent.
- Understanding of firmware, BIOS, BMC/IPMI/Redfish, and low-level system tuning.
- Proficiency in one or more of Go, Bash, Rust, or Python.
- Experience with GPU systems, RDMA fabrics, distributed AI workloads, virtualisation, bare-metal infrastructure, and Kubernetes/Slurm platforms.
- Strong debugging, problem-solving, documentation, and technical communication skills.
- Experience participating in an on-call rotation supporting production services.
Benefits
- Full-time employment.
- Work location in Singapore or Australia, including Melbourne, Sydney, or Launceston.
- Opportunity to work on sustainable AI infrastructure and large-scale GPU and storage platforms.