Graphcore

Staff Software Engineer - ML Kernels & Runtime

Graphcore
Apply
5 months ago
Gdańsk, PolandStaff+

Responsibilities

  • Design and implement C++ kernels for linear algebra and tensor operations including GEMM, batched GEMM, convolutions, reductions, elementwise operations, and fused operations.
  • Own kernel performance and correctness by adding microbenchmarks, regression tests, and numerical validation.
  • Profile and optimize threading, cache locality, memory layout, and kernel launch efficiency for next-generation AI hardware.
  • Debug issues, resolve bugs, and improve product quality and functionality.
  • Support Agile ways of working within the team.
  • Mentor colleagues by sharing knowledge and providing guidance.

Requirements

  • Excellent programming and scripting skills in C++ and Python.
  • Understanding of processor architectures and profiling on Linux.
  • Experience testing numerical and performance-sensitive code.
  • Hands-on experience with reproducibility, determinism, tolerance design, and benchmarking.
  • Strong written and oral communication, teamwork, and quality focus.
  • Strong algorithmic performance knowledge, including vectorization, memory hierarchy, threading, and lock-free patterns, is desirable.
  • Experience with at least one BLAS/DNN stack and extending kernels is desirable.
  • Experience with CPU micro-optimizations and numerical stability trade-offs across FP32, FP16, BF16, and FP8 is desirable.
  • Experience integrating native code into PyTorch or similar systems through custom operations, extensions, or dispatch keys is desirable.
  • Knowledge of ABI/API stability and Linux packaging, including manylinux and wheels, is desirable.

Benefits

  • Annual leave policy, medical and dental health plans, gym card, and employee pension matched up to 4%.
  • Flexible interview approach and reasonable adjustments are available.
Graphcore

About Graphcore

501-1,000 employees

Graphcore designs and sells Intelligence Processing Units (IPUs), systems, and a full software stack for training and inference of AI models in data centers and research labs. Revenue comes from hardware systems and associated software tools and services, used for workloads across NLP, computer vision, and graph neural networks. Founded in 2016 and headquartered in Bristol, the company is part of SoftBank Group and focuses on on‑prem and cloud deployments.

Contact me