Singtel

DevOps Engineer, GPUaaS

Singtel
Apply
1 month ago
Singapore, SingaporeMid Level

Responsibilities

  • Design, deploy, and support large-scale distributed GPU clusters for AI and machine learning workloads.
  • Manage and automate GPU resource provisioning across on-premises and cloud platforms.
  • Design, implement, and manage CI/CD pipelines for AI models and GPU-accelerated applications.
  • Monitor cluster usage, health, performance, and availability, and improve infrastructure management through automation.
  • Troubleshoot system-level issues involving Slurm, Kubernetes, GPU drivers, CUDA, and InfiniBand networking.
  • Optimize operating system, driver, networking, and library parameters for AI workload performance.
  • Benchmark GPU clusters and evaluate advances in GPU technology.
  • Set up GPU-resource monitoring and logging using Zabbix, Prometheus, NVIDIA DCGM, and related tools.
  • Implement security best practices for multi-tenant GPU-as-a-Service environments.
  • Provide technical support and guidance to users of GPU-accelerated systems.
  • Identify bottlenecks and improve development and operational processes for AI and HPC GPU cloud platforms.
  • Participate in rotational or scheduled shift work to support platform operations.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, Systems Engineering, or a related field.
  • Strong Linux system administration skills with Ubuntu, CentOS, Rocky Linux, or similar distributions.
  • Experience with Jenkins, Kubernetes, Ansible, and Terraform.
  • Understanding of DevOps practices including automation, monitoring, and CI/CD.
  • Proficiency in Python and Bash or similar scripting languages.
  • Experience implementing monitoring solutions such as Zabbix and Prometheus.
  • Familiarity with TensorFlow and PyTorch.
  • Understanding of IaaS and PaaS cloud architectures, GPU architecture, and NVIDIA GPUs.
  • Understanding of MPI, RDMA, NCCL, GPU acceleration, and HPC workload managers such as Slurm is desirable.
  • Knowledge of Docker or container technologies, data center deployments, InfiniBand, RoCE, DPUs, and NVIDIA GPU SDKs is desirable.
  • Strong English verbal, written, and presentation skills, with cross-functional coordination, analytical, and technical problem-solving abilities.

Tech Stack

AnsibleBashDockerJenkinsKubernetesLinuxPrometheusPythonPyTorchTensorFlowTerraform

Categories

Singtel

About Singtel

5,001-10,000 employees

Singtel is a Singapore‑headquartered public telecommunications group founded in 1879 that sells mobile, broadband, and enterprise network services to consumers and businesses across Asia-Pacific. It owns Optus in Australia and operates NCS (regional IT services) and Nxera (data centers), and is expanding into AI infrastructure via its RE:AI GPU‑as‑a‑Service. Its revenue comes from connectivity subscriptions and managed services across networks, cloud, cybersecurity, and digital platforms, with equity stakes in regional operators broadening its reach.

Contact me