21 days ago
Toronto, CanadaMid Level / Senior
Responsibilities
- Design, build, and maintain infrastructure, automation, and tooling to improve system reliability, availability, scalability, and operational efficiency.
- Partner with engineering teams to embed reliability, observability, and performance best practices throughout the software development lifecycle.
- Monitor production systems, troubleshoot and resolve issues, and minimize customer impact during service disruptions.
- Develop observability capabilities including metrics, logging, tracing, dashboards, and alerting.
- Participate in incident response and on-call support for critical systems and services.
- Conduct incident root cause analyses and drive corrective and preventative actions.
- Define and improve Service Level Objectives, Service Level Indicators, and operational performance metrics.
- Support disaster recovery, resilience testing, and business continuity initiatives.
- Promote proactive monitoring, automation, and operational excellence to reduce manual intervention.
- Collaborate with technical and business partners to improve service performance, stability, and customer experience.
Requirements
- Require 2–5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Operations, or a related technology field.
- Experience supporting applications and services in cloud environments such as Azure, AWS, or Google Cloud Platform.
- Experience developing automation and scripting solutions using Python, Bash, PowerShell, or similar technologies.
- Experience with monitoring, logging, and observability platforms such as New Relic, Grafana, Azure Data Explorer, or comparable tools.
- Strong troubleshooting, analytical, and problem-solving skills, including the ability to respond during service disruptions.
- Excellent communication and collaboration skills across technical and business teams.
- Preferred experience with Docker and Kubernetes.
- Preferred familiarity with ITIL practices and Agile delivery methodologies.
- Preferred experience creating operational dashboards and reporting using Power BI or similar analytics tools.
- Experience in financial services, banking, or other highly regulated environments is preferred.
Benefits
- Hybrid working arrangement in Waterloo, Ontario, as part of a distributed team.
- Flexible environment supporting learning, career growth, well-being, and inclusion.
- Eligible employees may receive health, dental, mental health, vision, disability, life, AD&D, adoption/surrogacy, wellness, and employee/family assistance benefits.
- Retirement savings plans may include pension and a global share ownership plan with employer matching contributions.
- Financial education and counseling resources are available.
- Paid time off in Canada includes holidays, vacation, personal, and sick days, plus statutory leaves of absence.
- Employees may participate in incentive programs tied to business and individual performance.
Tech Stack
Categories
DevOpsSite Reliability
