
Staff Platform Site Reliability Engineer
Index Exchange12 days ago
Toronto, CanadaStaff+
Responsibilities
- Design and deliver multi-tenant Kubernetes infrastructure across bare-metal and public cloud environments.
- Build infrastructure-as-code frameworks capable of deploying changes across thousands of servers.
- Develop standard libraries, SDKs, platform APIs, golden paths, and self-service tooling for engineering teams.
- Solve globally distributed systems challenges involving low-latency real-time bidding, multi-datacenter consistency, fleet deployment, and load balancing.
- Drive architectural direction through RFCs, design reviews, shared platform standards, tooling decisions, security posture, and system design.
- Mentor engineers, raise engineering standards, and collaborate across Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams.
Requirements
- 8+ years of experience in platform engineering, SRE, infrastructure engineering, or DevOps.
- Deep experience with Linux internals, including kernel tuning, networking, observability, and security.
- Strong Kubernetes expertise covering cluster lifecycle, networking, storage, RBAC, and multi-cluster environments across bare-metal and cloud.
- Experience building infrastructure-as-code at scale with Terraform, Ansible, and GitOps tools such as ArgoCD.
- Proficiency in Go, Python, or both for building libraries, SDKs, and platform APIs.
- Solid L2-L7 networking fundamentals, load balancing, DNS, and service discovery knowledge.
- Demonstrated ability to drive technical strategy across multiple teams.
- Valuable experience includes Ceph, Hadoop, Spark, HBase, Kafka, Prometheus, Grafana, ELK, Mimir, Loki, Tempo, Vault, certificate management, access control, hybrid cloud, and globally distributed bare-metal infrastructure.
Benefits
- Comprehensive health, dental, and vision plans for employees and dependents.
- Paid time off, health days, personal obligation days, and flexible work schedules.
- Competitive retirement matching plans, equity packages, and generous parental leave.
- Annual well-being allowance, fitness discounts, group wellness activities, commuter benefits where available, and an employee assistance program.
- Mental health first aid support, one volunteer day per year, donation matching, town halls, and community-led team events.
- Resources and programming for continuous learning, within a diverse, equitable, and inclusive workplace.
- The company supports talent across a global team with offices in Toronto, New York, Montreal, Kitchener, London, San Francisco, and other cities.
Tech Stack
AnsibleApache HadoopApache HBaseApache KafkaApache SparkAWSGoGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonTerraformVault
Categories
DevOpsSite Reliability