9 months ago
Toronto, CanadaSenior
Responsibilities
- Embed with product teams to improve service observability, reliability, performance, and scalability.
- Own CI/CD pipelines, observability tooling, monitoring systems, and incident response processes.
- Build production-grade software, tools, and automation to reduce toil and improve engineering velocity, developer experience, and system reliability.
- Collaborate with engineering teams at the code level to identify cross-cutting reliability, performance, and scaling concerns.
- Architect and scale infrastructure for performance, availability, and operational excellence.
- Drive capacity planning and define and manage SLOs and error budgets with service-owning engineering teams.
- Advocate for reliability, performance, and scalability across the engineering organization.
Requirements
- 5+ years of experience in an SRE, Platform, DevOps, or Infrastructure Engineering role.
- 5+ years of experience writing software in a production environment.
- Strong knowledge of cloud infrastructure, distributed systems, reliability practices, observability, performance tuning, and scaling strategies.
- Deep familiarity with incident response, monitoring, and CI/CD systems.
- Hands-on experience supporting web or RPC services at meaningful scale.
- Ability to write production-grade software for infrastructure problems rather than relying only on shell scripts.
- Systems-level mindset, proactive reliability approach, ownership of complex problems, and ability to influence product-team design and architecture decisions.
- Ruby and Go experience preferred.
Benefits
- Competitive compensation and early equity at a fast-growing, venture-backed company.
- Comprehensive medical, dental, and vision coverage.
- Three weeks of vacation, unlimited sick and mental health days, and a company-wide end-of-year shutdown.
- $500 home office setup stipend.
- Unlimited token usage and access to AI tools.
- High-impact environment with substantial ownership and leadership opportunities.
Tech Stack
Categories
DevOpsSite Reliability