G-Re

Machine Learning Performance Engineer

G-Re
Apply
2 hours ago
London, United KingdomSenior

Responsibilities

  • Collaborate with researchers, senior stakeholders, and engineers to understand compute challenges and design optimized solutions.
  • Profile, benchmark, and tune large-scale training and inference workloads across distributed CPU, GPU, and memory-intensive jobs.
  • Develop reference implementations, libraries, and tools that improve job efficiency and reliability.
  • Work with systems, architecture, and platform teams to evolve the compute stack.
  • Influence long-term platform and infrastructure decisions.

Requirements

  • Bachelor’s, Master’s, or PhD degree in computer science, or equivalent experience.
  • Proven track record profiling, benchmarking, and optimizing distributed workloads.
  • Experience with Python and knowledge of CUDA.
  • Experience with HPC schedulers and Kubernetes-based workload orchestration.
  • Strong understanding of a deep learning framework such as PyTorch.
  • Strong background in data structures, algorithms, and parallel programming on heterogeneous systems.
  • Deep understanding of Linux fundamentals, including scheduling, memory management, NUMA, networking, and filesystems.
  • Familiarity with nsys, ncu, eBPF-based tools, and performance counters.
  • Strong communication and cross-functional collaboration skills.

Benefits

  • Highly competitive compensation plus an annual discretionary bonus.
  • Lunch provided via Just Eat for Business and access to a dedicated barista bar.
  • 35 days’ annual leave.
  • 9% company pension contributions.
  • Informal dress code and work/life balance.
  • Comprehensive healthcare and life assurance.
  • Cycle-to-work scheme.
  • Monthly company events.
  • Accommodation support is available for applicants with a disability or special need.
  • The role is based from the company’s London headquarters.
G-Re

About G-Re

1,001-5,000 employees
Contact me