about 2 hours ago
Responsibilities
- Translate System Management direction into technical plans, engineering priorities, and deliverables.
- Lead architecture and design decisions, document trade-offs, and build alignment across teams.
- Act as the technical authority for assigned System Management areas and lead complex engineering initiatives.
- Coordinate technical plans, dependencies, risks, decisions, and delivery priorities.
- Provide technical direction and mentoring across system management, hardware lifecycle management, deployment automation, and production operations.
- Own technical outcomes across design, implementation, automated testing, integration, deployment, observability, and production readiness.
- Identify and improve systemic reliability, scalability, and operability issues.
- Collaborate with Hardware, Firmware, Platform Software, and Datacenter Operations teams on system-level issues.
- Improve CI/CD, Infrastructure-as-Code, automated testing, release safety, and operational learning practices.
- Serve as a senior technical escalation point and create reusable tooling, automation, and operational knowledge.
Requirements
- Bachelor’s degree in a relevant subject or equivalent practical experience.
- Substantial experience designing, building, and operating Linux-based infrastructure or distributed systems.
- Demonstrated technical leadership of complex, multi-team engineering initiatives and ability to influence architecture across team boundaries.
- Strong experience developing RESTful APIs and programming in Go, with Bash and Python for systems automation.
- Deep practical experience with Kubernetes, container runtimes, and production workloads.
- Hands-on experience with Infrastructure-as-Code, source control, and CI/CD technologies including Terraform/OpenTofu, Ansible, GitLab, GitHub Actions, and Git.
- Experience with hardware-management interfaces such as Redfish and IPMI.
- Strong Linux systems engineering, troubleshooting, and operational debugging skills.
- Experience mentoring and developing engineers through technical mentoring, design reviews, and coaching.
- Clear communication and stakeholder-alignment skills.
- Experience using AI coding assistants is desirable.
- Experience developing Kubernetes operators and custom resources is desirable.
- Experience with HPC environments using SLURM, LSF, or similar workload-management systems is desirable.
- Experience with Open vSwitch, KVM, and QEMU is desirable.
- Experience with distributed object, block, and file storage technologies such as Ceph is desirable.
- Experience with Grafana, Prometheus, OpenSearch/Elasticsearch, Loki, Mimir, or OpenTelemetry is desirable.
- Experience configuring managed network switches using EOS, SONiC, or DNOS is desirable.
- Experience supporting AI infrastructure or PyTorch workloads is desirable.
Benefits
- Flexible working arrangement.
- Medical, dental, and vision coverage.
- Flexible Spending Accounts and Health Savings Accounts.
- Disability and life insurance.
- 401(k) retirement plan.
- Commuter benefits, wellness services, and Employee Assistance Programme.
- Inclusive workplace and reasonable interview adjustments.
Tech Stack
About Graphcore
At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.