Graphcore

Principal Engineer

Graphcore
Apply
2 months ago
Austin, TX, USA or Milpitas, CA, USAStaff+
H1B sponsor

Responsibilities

  • Translate System Management direction into technical plans, priorities, milestones, and deliverables.
  • Lead architecture and design decisions while documenting trade-offs and building alignment across teams.
  • Lead complex engineering initiatives and coordinate dependencies, risks, decisions, and delivery plans.
  • Provide technical direction and mentoring across system management, hardware lifecycle management, deployment automation, and production operations.
  • Own technical outcomes across design, implementation, automated testing, integration, deployment, observability, and production readiness.
  • Identify and improve systemic reliability, scalability, and operability issues.
  • Collaborate with Hardware, Firmware, Platform Software, and Datacenter Operations teams on system-level issues.
  • Improve engineering practices involving Infrastructure-as-Code, automated testing, release safety, CI/CD, and operational learning.
  • Serve as a senior technical escalation point and create reusable tooling, automation, and knowledge to reduce operational effort.

Requirements

  • Bachelor’s degree in a relevant subject or equivalent practical experience.
  • Substantial experience designing, building, and operating Linux-based infrastructure or distributed systems.
  • Experience providing technical leadership for complex, multi-team engineering initiatives and influencing architecture across team boundaries.
  • Strong experience with Go, Bash, Python, and RESTful APIs.
  • Deep practical experience with Kubernetes, container runtimes, and production workloads.
  • Hands-on experience with Infrastructure-as-Code, source control, and tools including Terraform/OpenTofu, Ansible, GitLab, GitHub Actions, and Git.
  • Experience with hardware-management interfaces such as Redfish and IPMI or equivalent systems.
  • Strong Linux systems engineering, troubleshooting, and operational debugging skills.
  • Experience mentoring and developing engineers through design reviews and coaching.
  • Experience with Kubernetes operators and custom resources is desirable.
  • Experience with HPC workload-management systems such as SLURM or LSF is desirable.
  • Experience with Open vSwitch, KVM, QEMU, Ceph, observability platforms, managed network switches, or AI infrastructure and PyTorch is desirable.

Benefits

  • Flexible working arrangement
  • Medical, dental and vision coverage
  • Flexible Spending Accounts (FSAs)
  • Health Savings Accounts (HSAs)
  • Disability and life insurance
  • 401(k) retirement plan
  • Commuter benefits
  • Wellness services
  • Employee Assistance Programme (EAP)
  • Inclusive equal-opportunity workplace with reasonable interview adjustments

Tech Stack

AnsibleBashElasticsearchGitGitHub ActionsGoGrafanaKubernetesLinuxPrometheusPythonPyTorchTerraform

Categories

Graphcore

About Graphcore

501-1,000 employees

Graphcore designs and sells Intelligence Processing Units (IPUs), systems, and a full software stack for training and inference of AI models in data centers and research labs. Revenue comes from hardware systems and associated software tools and services, used for workloads across NLP, computer vision, and graph neural networks. Founded in 2016 and headquartered in Bristol, the company is part of SoftBank Group and focuses on on‑prem and cloud deployments.

Contact me