3 months ago
Responsibilities
- Own software engineering efforts across the full software development lifecycle, including implementation, automated testing, integration, and production readiness for the rack-management solution.
- Own critical infrastructure and drive issues through resolution while collaborating across teams.
- Configure and test Graphcore AI hardware and systems using Continuous Deployment and Infrastructure-as-Code in internal and external datacentres.
- Work with Datacenter Operations Engineers to maintain and operate the fleet of AI systems at peak performance.
- Drive corrective actions for systems that are not operating correctly in partnership with datacenter operations and Graphcore Engineering.
- Provide hands-on support and troubleshooting across complex system-level environments.
Requirements
- Bachelor’s degree or equivalent practical experience in a relevant subject.
- Experience developing RESTful APIs.
- Experience building, deploying, and operating containerized workloads with Kubernetes and Docker or Podman.
- Experience managing production Kubernetes clusters and workloads.
- Programming experience with Go.
- Hands-on experience deploying and operating infrastructure with Infrastructure-as-Code, source control, and CI/CD tools such as Terraform/OpenTofu, Ansible, GitLab, GitHub Actions, and Git.
- Experience with Redfish for datacenter hardware management, telemetry, provisioning, and control.
- Experience specifying, scoping, estimating, and detailing work plans within Agile and Scrum frameworks.
- Strong Linux systems engineering experience, including administration, automation, and Bash and Python scripting.
- Desirable experience with AI coding assistants such as Codex and Claude.
- Desirable experience developing Kubernetes operators and custom resources.
- Desirable experience with HPC environments using SLURM or similar batch workload solutions.
- Desirable experience with virtualized deployments and Open vSwitch, KVM, or QEMU.
- Desirable experience with distributed object, block, and file storage such as Ceph.
- Desirable experience with end-to-end deployment automation and CI for containerized services.
- Desirable experience with monitoring and observability tools such as Grafana, Prometheus, OpenSearch, Elasticsearch, Loki, Mimir, OpenTelemetry, Fluentd, and Kafka.
- Desirable experience configuring managed switches using EOS, SONiC, or DNOS.
- Desirable experience with PyTorch for AI workloads.
- Understanding of cloud and infrastructure technologies including APIs, virtualization, networking, block storage, resource management, and monitoring systems.
Tech Stack
AnsibleApache KafkaBashDockerElasticsearchGitGitHub ActionsGoGrafanaKubernetesLinuxPrometheusPythonPyTorchTerraform
About Graphcore
At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.