
Staff Platform Site Reliability Engineer
Index Exchange12 days ago
London, United KingdomStaff+
Responsibilities
- Design and deliver multi-tenant Kubernetes infrastructure across bare-metal and public cloud environments.
- Build infrastructure-as-code frameworks that deploy changes across thousands of servers.
- Develop standard libraries, internal SDKs, platform APIs, golden paths, and self-service tooling.
- Solve distributed-systems problems involving real-time bidding, multi-datacenter consistency, fleet deployment, and large-scale load balancing.
- Own architecture through RFCs and design reviews and set standards for shared platform domains.
- Drive technical direction across multiple divisions and make decisions about tooling, security posture, and system design.
- Mentor engineers and collaborate with Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams.
Requirements
- 8+ years of experience in platform engineering, SRE, infrastructure engineering, or DevOps.
- Deep experience with Linux internals, including kernel tuning, network stack, system observability, and security.
- Strong Kubernetes expertise covering cluster lifecycle, networking, storage, RBAC, and multi-cluster environments across bare-metal and cloud.
- Experience with infrastructure-as-code at scale using Terraform, Ansible, and GitOps such as ArgoCD.
- Proficiency in Go, Python, or both for building libraries, SDKs, and platform APIs.
- Strong networking fundamentals across L2-L7, including load balancing, DNS, and service discovery.
- Track record of driving technical strategy across teams.
- Valuable experience includes Ceph; Hadoop, Spark, HBase, or Kafka; Prometheus, Grafana, ELK, Mimir, Loki, or Tempo; Vault; certificate management; access control; hybrid cloud architectures; and globally distributed bare-metal infrastructure.
- Ability to understand complex distributed systems, prioritize operational readiness, build effective abstractions, align teams, and improve developer productivity.
Benefits
- Comprehensive health, dental, and vision plans for employees and dependents.
- Paid time off, health days, personal obligation days, and flexible work schedules.
- Competitive retirement matching plans and equity packages.
- Generous parental leave for birthing, non-birthing, and adoptive parents.
- Annual well-being allowance, fitness discounts, and group wellness activities.
- Commuter benefits and discounts where available.
- Employee assistance program and mental health first aid program.
- One day of volunteer time off annually and a donation-matching program.
- Monthly town halls, community-led team events, and continuous-learning resources.
- Global workplace with offices in Toronto, New York, Montreal, Kitchener, London, San Francisco, and other cities.
Tech Stack
AnsibleApache HadoopApache HBaseApache KafkaApache SparkAWSGoGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonTerraformVault
Categories
DevOpsSite Reliability