
Industrial AI Cloud - Platform Engineer (REF5501C)
Deutsche Telekom IT Solutions4 months ago
Budapest, HungarySenior
Responsibilities
- Build, configure, maintain, and operate bare-metal hosts and large-scale Kubernetes clusters for GPU and AI workloads.
- Design and operate the NVIDIA AI software stack, including Slurm and Run:ai, and provide application support for AI workloads.
- Manage Helm charts, GitOps workflows, Ansible scripts, Terraform automation, and Kubernetes resources.
- Develop and maintain Jenkins and GitLab CI/CD pipelines and implement consistent deployment and infrastructure-change practices.
- Troubleshoot, performance-tune, scale, and serve as the primary contact for Kubernetes-related topics.
- Operate Prometheus and Grafana monitoring and observability stacks for hosts, Kubernetes, and platform services.
- Build, optimize, secure, scan, and manage Docker and Podman container images and registries.
- Integrate and maintain object storage and persistent-volume solutions for AI workloads.
- Support distributed AI and HPC workloads across bare-metal and Kubernetes environments.
- Collaborate with infrastructure engineers, data center staff, security teams, and AI developers to deliver reliable services.
- Follow incident, problem, and change-management workflows, maintain operational runbooks, and adhere to zero-outage guidelines.
Requirements
- Production Kubernetes experience, with CKA certification or equivalent experience required and CKS considered an advantage.
- Knowledge of NVIDIA GPU-accelerated server platforms, NVIDIA AI software, GPU orchestration, and GPU-based cloud-platform dependencies.
- Knowledge of data engineering, data transformation, and data migration tools.
- Strong experience with Jenkins, GitLab, CI/CD tooling, and GitOps practices.
- Proficiency with Helm charts and Kubernetes resource management.
- Ability to script or program in Python or Bash.
- Experience with Infrastructure as Code using Terraform and Ansible.
- Experience with Docker, Podman, container registries, and image-scanning tools.
- Familiarity with object storage, persistent-volume management, Prometheus, and Grafana.
- Understanding of AI and HPC workloads at scale and strong troubleshooting and operational-support skills in mission-critical environments.
- Ability to automate repetitive operational tasks and build self-service capabilities for developers.
- Security-conscious working approach and strong communication and cross-team coordination skills.
Benefits
- Work on Europe’s first industrial AI cloud using cutting-edge technologies.
- Collaborate directly with NVIDIA and Deutsche Telekom experts.
- Hybrid working model; remote work is available only within Hungary due to European taxation regulations.
- Training opportunities and career progression.
- The role is based in the European Union to meet customer data-security and privacy requirements.