15 hours ago
Responsibilities
- Analyze and optimize AI training and inference workloads across distributed systems.
- Lead complex performance workstreams spanning hardware and software.
- Identify bottlenecks, validate improvements, and turn performance evidence into engineering decisions.
- Build benchmarks, profile demanding workloads, and develop C++ and Python tools for analysis and optimization.
- Coordinate cross-team investigations and validate impact using reliable performance data.
- Improve the efficiency, scalability, and reliability of large-scale AI infrastructure.
Requirements
- Strong experience profiling and optimizing AI, machine learning, or high-performance computing workloads.
- Experience with distributed systems and communication libraries such as MPI, NCCL, UCX, or libfabric.
- Strong C++ and Python skills, including experience building reliable tools or performance-sensitive software.
- Deep understanding of compute, memory, and communication behavior in large-scale systems.
- Ability to lead technically complex work and coordinate improvements across teams.
- Familiarity with MLPerf, accelerated architectures, ML frameworks, or high-performance interconnects.
Benefits
- Medical, dental, and vision coverage, with options that may extend to eligible dependents.
- Mental health, wellness, and employee assistance resources.
- Retirement savings benefits and company contributions where applicable.
- Paid vacation, sick time, company holidays, and parental or family leave.
- Life insurance and short-term or long-term disability coverage.
- Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
- Professional-development resources, learning programs, office amenities, and team-led activities.
- The role is based in Milpitas, California.
Categories
About Graphcore
Graphcore designs and sells Intelligence Processing Units (IPUs), systems, and a full software stack for training and inference of AI models in data centers and research labs. Revenue comes from hardware systems and associated software tools and services, used for workloads across NLP, computer vision, and graph neural networks. Founded in 2016 and headquartered in Bristol, the company is part of SoftBank Group and focuses on on‑prem and cloud deployments.