Senior AI Engineer (Kubernetes & Customised Scheduler)
Firmus Technologies4 hours ago
Singapore, SingaporeSenior
Responsibilities
- Design, build, operate, and continuously improve a proprietary Kubernetes-native job scheduling platform for AI workloads.
- Develop custom resource definitions, Kubernetes controllers, admission webhooks, scheduling plugins, APIs, CLI tools, and automation for workload submission, placement, execution, monitoring, and recovery.
- Implement scheduling policies covering training, fine-tuning, inference, benchmarking, batch, data-processing, and agentic workloads.
- Build topology-aware and AI-factory-resource-aware placement mechanisms using GPU, NVLink/NVSwitch, NUMA, NIC, RDMA, storage, capacity, health, power, thermal, maintenance, and fault-domain signals.
- Implement admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery.
- Integrate Kubernetes with GPU device plugins, node-feature discovery, network and storage services, observability systems, policy engines, and workload-management systems.
- Create APIs, SDKs, CLI workflows, templates, workload definitions, status visibility, event streams, scheduling explanations, and self-service troubleshooting capabilities.
- Expose scheduler, placement, resource, and workload-lifecycle data for Model-to-Grid benchmarking and performance analysis.
- Build observability for queue depth, latency, placement decisions, resource fragmentation, utilization, failures, retries, preemption, power, thermal signals, and workload performance.
- Operate production Kubernetes clusters and scheduling services using reliability, security, change-management, backup, recovery, and incident-response practices.
- Build CI/CD and GitOps workflows for scheduler code, configuration, CRDs, controllers, policies, templates, and cluster-service releases.
- Implement secure multi-tenancy and governance controls including tenant isolation, RBAC, quotas, workload identity, policy enforcement, audit trails, and controlled resource access.
- Partner with product, platform, infrastructure, networking, storage, security, and operations teams on architecture, milestones, releases, benchmarks, documentation, and continuous improvement.
Requirements
- At least 5 years of experience in DevOps, site reliability engineering, platform engineering, distributed systems, cloud infrastructure, or comparable software and systems engineering roles.
- Deep practical expertise in Kubernetes architecture and operations, including control planes, scheduling, controllers, operators, CRDs, admission controllers, APIs, networking, storage, security, and multi-tenancy.
- Demonstrated experience building, extending, or deeply integrating workload schedulers, resource managers, or orchestration systems such as Kubernetes Scheduler Framework, Kueue, KAI, Volcano, YuniKorn, Slurm, Slinky, Run:ai, or proprietary implementations.
- Strong Go programming skills and Python proficiency for automation, integration, tooling, and performance analysis.
- Experience building production Kubernetes controllers, operators, scheduling plugins, admission webhooks, REST or gRPC APIs, CLI tools, and event-driven distributed services.
- Strong understanding of GPU-accelerated AI/ML workloads, distributed training, fine-tuning, model serving, inference, benchmarking, and batch processing.
- Experience with GPU scheduling and resource allocation, including GPU partitioning, MIG, GPU affinity, multi-GPU workloads, gang scheduling, and topology-aware placement.
- Understanding of GPU system topology and high-performance AI networking, including NVLink, NVSwitch, PCIe, NUMA, RDMA, RoCEv2, NIC affinity, collective communication, and network contention.
- Familiarity with large-scale NVIDIA GPU architectures, including NVL72 GB300-class systems and future-generation high-density GPU platforms.
- Experience with AI workload storage and data paths, including shared file systems, object storage, caching, data locality, checkpointing, and storage-performance constraints.
- Understanding of AI-factory or data-center operations including capacity planning, demand forecasting, power-aware scheduling, energy efficiency, thermal constraints, maintenance coordination, resiliency, and infrastructure telemetry.
- Experience with observability and performance analysis using Prometheus, Grafana, OpenTelemetry, logs, traces, metrics, DCGM, GPU telemetry, network telemetry, workload profiling, and SLOs.
- Experience with CI/CD, GitOps, infrastructure as code, policy as code, and safe production rollout practices using GitHub Actions, GitLab CI, Argo CD, Flux, Terraform, Helm, Kustomize, OPA, or Kyverno.
- Familiarity with cloud-native security practices including container security, image signing, software supply-chain controls, secrets management, RBAC, workload identity, network policies, runtime security, and vulnerability remediation.
- Strong troubleshooting skills across Kubernetes, distributed systems, GPU workloads, networking, storage, operating systems, and application-level job execution.
Benefits
- Permanent full-time employment.
- Location options in Singapore or Australia, including Launceston, Hobart, Sydney, and Melbourne.
- Opportunity to work with founders and experts in AI infrastructure, energy systems, and next-generation compute.
- Exposure to AI infrastructure as an NVIDIA Cloud and Engineering partner in Asia Pacific.
- Work on sustainable AI-factory infrastructure designed to improve energy efficiency and support surrounding communities.