over 3 years ago
Toronto, CanadaSenior
Responsibilities
- Embed with product teams to improve service observability, reliability, performance, and scalability.
- Own CI/CD pipelines, observability tooling, monitoring systems, and incident response processes.
- Build production-grade software, tools, and automation to reduce manual toil and improve developer experience and system reliability.
- Architect and scale infrastructure while driving capacity planning and operational excellence.
- Define and manage SLOs and error budgets with engineering teams responsible for production services.
- Identify and address cross-cutting reliability, performance, and scaling concerns at the code and systems level.
Requirements
- Require 5+ years of experience in an SRE, platform, or infrastructure engineering role.
- Require 5+ years of experience writing software in a production environment.
- Strong knowledge of cloud infrastructure, distributed systems, reliability practices, observability, performance tuning, and scaling strategies.
- Deep familiarity with incident response, monitoring, and CI/CD systems.
- Hands-on experience supporting web or RPC services at meaningful scale.
- Ability to write production-grade software to solve infrastructure problems rather than relying on shell scripts alone.
- Preferred: experience embedding with product teams and influencing design and architecture decisions.
- Preferred: experience with Ruby and Go, a systems-level mindset, and ownership of complex problems.
Benefits
- Competitive compensation and early equity at a venture-backed company.
- Comprehensive medical, dental, and vision coverage.
- Three weeks of vacation, unlimited sick and mental health days, and a company-wide end-of-year shutdown.
- Home office setup stipend.
- Unlimited token usage and access to AI tools.
- High-impact, fast-moving environment with significant ownership and leadership influence.
Tech Stack
GoRuby
Categories
Site Reliability