3 months ago
Gdańsk, PolandSenior
Responsibilities
- Implement current machine learning models and optimize them for performance and accuracy across thousands of accelerators.
- Test and evaluate internal software releases, provide feedback, make code fixes, and conduct code reviews.
- Benchmark models and ML techniques to identify performance bottlenecks and improve efficiency.
- Design, implement, and evaluate experiments involving novel AI methods.
- Collaborate with Research, Software, and Product teams to define, build, and test next-generation AI hardware.
- Engage with the AI community and track the latest developments in artificial intelligence.
Requirements
- Bachelor’s, Master’s, PhD, or equivalent experience in Machine Learning, Computer Science, Mathematics, Data Science, or a related field.
- Proficiency with deep learning frameworks such as PyTorch and JAX.
- Strong Python or C++ software development skills.
- Expertise in deep learning, including model training, optimization, and evaluation.
- Experience with distributed ML model training or inference across 64 or more accelerators.
- Ability to design, execute, and report on ML experiments.
- Deep understanding of performance bottlenecks and methods to overcome them.
- Desirable experience with Kubernetes-based MLOps, production systems using large language models, or low-precision arithmetic.
- Desirable experience writing C++, Triton, or CUDA kernels for ML model performance optimization.
- Familiarity with HPC systems and networking technologies including InfiniBand, NVLink, and RoCE.
- Open-source contributions or relevant research publications are desirable.
- Knowledge of cloud computing platforms is desirable.
- Strong communication and cross-functional collaboration skills.
Benefits
- Annual leave policy
- Medical and dental health plans
- Gym card
- Employee pension matched up to 4%
- Flexible interview approach and reasonable-adjustment support
Tech Stack
Categories
About Graphcore
Graphcore designs and sells Intelligence Processing Units (IPUs), systems, and a full software stack for training and inference of AI models in data centers and research labs. Revenue comes from hardware systems and associated software tools and services, used for workloads across NLP, computer vision, and graph neural networks. Founded in 2016 and headquartered in Bristol, the company is part of SoftBank Group and focuses on on‑prem and cloud deployments.