
Senior ML Infrastructure Engineer
Ellison Institute of Technology9 months ago
Oxford, United KingdomSenior
Responsibilities
- Build, operate, and continuously optimise high-performance GPU training and inference clusters.
- Design and implement high-throughput data paths, including I/O, caching, and data locality across compute and storage.
- Benchmark, profile, and resolve performance bottlenecks across compute, network, and orchestration layers.
- Establish observability, resilience, and automated security controls for sensitive research environments.
- Partner with Research, Data, and Applied teams on GPU and storage capacity and cost forecasting, quotas, and ML experimentation pipelines.
Requirements
- Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale.
- Ability to independently design, ideate, collaborate on, and implement effective systems solutions.
- Experience migrating or transforming ML infrastructure from traditional schedulers to modern containerised systems.
- Expertise with high-throughput storage systems for ML/HPC workloads.
- Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling.
- Knowledge of infrastructure-as-code and CI/CD practices, including Terraform and Argo CD.
Benefits
- Enhanced holiday pay, pension, life assurance, income protection, private medical insurance, hospital cash plan, therapy services, Perk Box, and an electric car scheme.
- Hybrid workplace in Oxford, United Kingdom.
Tech Stack
Argo CDTerraform