
Senior GPU Systems & Fabric Engineer
BitDeer Technologies Group1 month ago
Singapore, SingaporeSenior
Responsibilities
- Architect and maintain NVIDIA/AMD GPU device plugins and Kubernetes Operators for exposing hardware capabilities to the control plane.
- Configure and optimize RDMA, SR-IOV, RoCEv2, and InfiniBand networking for distributed AI training.
- Build automated DCGM-based remediation pipelines to identify, isolate, and reset degraded GPU and NIC components.
- Implement MIG and vGPU slicing for efficient multi-tenant inference workloads.
- Profile and tune kernel parameters, device drivers, and CUDA/NCCL runtime libraries for containerized AI workloads.
- Collaborate with scheduling and storage teams on topology-aware placement and efficient data movement.
- Define operational standards for bare-metal provisioning, BIOS and firmware updates, and OS hardening.
- Lead investigations into performance issues across hardware, fabric, and software; provide architectural recommendations.
- Mentor team members and drive documentation standards for the AI hardware stack.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 5+ years of systems engineering experience.
- Strong proficiency with Linux kernel internals and C or Go.
- Hands-on experience with NVIDIA H100/A100 GPU architectures, CUDA runtimes, RDMA, and InfiniBand.
- Deep understanding of containerized environments and Kubernetes device plugin architecture.
- Experience operating, debugging, and scaling bare-metal systems in large-scale production or HPC environments.
- Familiarity with Terraform, Ansible, and CI/CD pipelines for hardware lifecycle management.
- Strong problem-solving and communication skills with a collaborative approach across infrastructure, scheduling, and reliability teams.
- Experience in high-velocity, high-growth engineering environments is strongly preferred.
Benefits
- Inclusive and diverse work environment with open workspaces and a start-up atmosphere.
- Opportunity to contribute to digital asset and AI infrastructure projects and work with industry pioneers.
- Personal accountability, autonomy, professional growth, and learning opportunities.
- Training, mentoring, welfare benefits, and developmental opportunities.