15 hours ago
Responsibilities
- Define performance direction for large-scale AI training and inference workloads.
- Lead performance investigations spanning model behavior, profiling, benchmarking, modeling, and multi-node scaling.
- Turn performance evidence into technical roadmaps and engineering priorities.
- Define benchmarks, measurement practices, and optimization standards across teams.
- Guide distributed communication strategy and review critical C++ and Python tools.
- Influence hardware, software, networking, and system-architecture decisions across the organization.
Requirements
- Deep experience profiling and optimizing AI, machine learning, or high-performance computing workloads at scale.
- Proven technical leadership across complex performance-engineering or system-architecture initiatives.
- Expert C++ and Python skills, including development of reliable tools or performance-sensitive software.
- Expert understanding of compute, memory, communication, and scaling behavior in distributed systems.
- Experience defining benchmarks, measurement practices, or optimization standards across teams.
- Experience with MPI, NCCL, UCX, libfabric, MLPerf, accelerated architectures, or high-performance interconnects.
- Transferable skills and diverse professional experiences are welcomed, including engineers returning after a career break through returnship routes.
Benefits
- Medical, dental, and vision coverage, potentially including eligible dependents.
- Mental health, wellness, and employee assistance resources.
- Retirement savings benefits and company contributions where applicable.
- Paid vacation, sick time, company holidays, and parental or family leave according to applicable plans and policies.
- Life insurance and short- and long-term disability coverage.
- Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
- Professional-development resources, learning programs, office amenities, and team-led activities.
- The role is based in Austin, Texas, with benefits varying by work location, employment status, scheduled hours, and plan eligibility.
Categories
About Graphcore
Graphcore designs and sells Intelligence Processing Units (IPUs), systems, and a full software stack for training and inference of AI models in data centers and research labs. Revenue comes from hardware systems and associated software tools and services, used for workloads across NLP, computer vision, and graph neural networks. Founded in 2016 and headquartered in Bristol, the company is part of SoftBank Group and focuses on on‑prem and cloud deployments.