1 month ago
Singapore, SingaporeMid Level
Responsibilities
- Design, deploy, and support large-scale distributed GPU clusters for AI and machine learning workloads.
- Manage and automate GPU resource provisioning across on-premises and cloud platforms.
- Design, implement, and manage CI/CD pipelines for AI models and GPU-accelerated applications.
- Monitor cluster usage, health, performance, and availability, and improve infrastructure management through automation.
- Troubleshoot system-level issues involving Slurm, Kubernetes, GPU drivers, CUDA, and InfiniBand networking.
- Optimize operating system, driver, networking, and library parameters for AI workload performance.
- Benchmark GPU clusters and evaluate advances in GPU technology.
- Set up GPU-resource monitoring and logging using Zabbix, Prometheus, NVIDIA DCGM, and related tools.
- Implement security best practices for multi-tenant GPU-as-a-Service environments.
- Provide technical support and guidance to users of GPU-accelerated systems.
- Identify bottlenecks and improve development and operational processes for AI and HPC GPU cloud platforms.
- Participate in rotational or scheduled shift work to support platform operations.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, Systems Engineering, or a related field.
- Strong Linux system administration skills with Ubuntu, CentOS, Rocky Linux, or similar distributions.
- Experience with Jenkins, Kubernetes, Ansible, and Terraform.
- Understanding of DevOps practices including automation, monitoring, and CI/CD.
- Proficiency in Python and Bash or similar scripting languages.
- Experience implementing monitoring solutions such as Zabbix and Prometheus.
- Familiarity with TensorFlow and PyTorch.
- Understanding of IaaS and PaaS cloud architectures, GPU architecture, and NVIDIA GPUs.
- Understanding of MPI, RDMA, NCCL, GPU acceleration, and HPC workload managers such as Slurm is desirable.
- Knowledge of Docker or container technologies, data center deployments, InfiniBand, RoCE, DPUs, and NVIDIA GPU SDKs is desirable.
- Strong English verbal, written, and presentation skills, with cross-functional coordination, analytical, and technical problem-solving abilities.
Categories
About Singtel
Singtel is a Singapore‑headquartered public telecommunications group founded in 1879 that sells mobile, broadband, and enterprise network services to consumers and businesses across Asia-Pacific. It owns Optus in Australia and operates NCS (regional IT services) and Nxera (data centers), and is expanding into AI infrastructure via its RE:AI GPU‑as‑a‑Service. Its revenue comes from connectivity subscriptions and managed services across networks, cloud, cybersecurity, and digital platforms, with equity stakes in regional operators broadening its reach.
