Ellison Institute of Technology

Senior ML Infrastructure Engineer

Ellison Institute of Technology
Apply
9 months ago
Oxford, United KingdomSenior

Responsibilities

  • Build, operate, and continuously optimise high-performance GPU training and inference clusters.
  • Design and implement high-throughput data paths, including I/O, caching, and data locality across compute and storage.
  • Benchmark, profile, and resolve performance bottlenecks across compute, network, and orchestration layers.
  • Establish observability, resilience, and automated security controls for sensitive research environments.
  • Partner with Research, Data, and Applied teams on GPU and storage capacity and cost forecasting, quotas, and ML experimentation pipelines.

Requirements

  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale.
  • Ability to independently design, ideate, collaborate on, and implement effective systems solutions.
  • Experience migrating or transforming ML infrastructure from traditional schedulers to modern containerised systems.
  • Expertise with high-throughput storage systems for ML/HPC workloads.
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling.
  • Knowledge of infrastructure-as-code and CI/CD practices, including Terraform and Argo CD.

Benefits

  • Enhanced holiday pay, pension, life assurance, income protection, private medical insurance, hospital cash plan, therapy services, Perk Box, and an electric car scheme.
  • Hybrid workplace in Oxford, United Kingdom.

Tech Stack

Argo CDTerraform
Ellison Institute of Technology

About Ellison Institute of Technology

201-500 employees
Contact me