2 hours ago
London, United KingdomSenior
Responsibilities
- Collaborate with researchers, senior stakeholders, and engineers to understand compute challenges and design optimized solutions.
- Profile, benchmark, and tune large-scale training and inference workloads across distributed CPU, GPU, and memory-intensive jobs.
- Develop reference implementations, libraries, and tools that improve job efficiency and reliability.
- Work with systems, architecture, and platform teams to evolve the compute stack.
- Influence long-term platform and infrastructure decisions.
Requirements
- Bachelor’s, Master’s, or PhD degree in computer science, or equivalent experience.
- Proven track record profiling, benchmarking, and optimizing distributed workloads.
- Experience with Python and knowledge of CUDA.
- Experience with HPC schedulers and Kubernetes-based workload orchestration.
- Strong understanding of a deep learning framework such as PyTorch.
- Strong background in data structures, algorithms, and parallel programming on heterogeneous systems.
- Deep understanding of Linux fundamentals, including scheduling, memory management, NUMA, networking, and filesystems.
- Familiarity with nsys, ncu, eBPF-based tools, and performance counters.
- Strong communication and cross-functional collaboration skills.
Benefits
- Highly competitive compensation plus an annual discretionary bonus.
- Lunch provided via Just Eat for Business and access to a dedicated barista bar.
- 35 days’ annual leave.
- 9% company pension contributions.
- Informal dress code and work/life balance.
- Comprehensive healthcare and life assurance.
- Cycle-to-work scheme.
- Monthly company events.
- Accommodation support is available for applicants with a disability or special need.
- The role is based from the company’s London headquarters.
