13 hours ago
Toronto, CanadaSenior
Responsibilities
- Eliminate operational toil through automation and software development.
- Build middleware, platform guardrails, orchestration, and internal SaaS tools that improve developer productivity and service reliability.
- Scale and maintain a resilient, cost-effective, secure-by-default cloud platform and Kubernetes ecosystem.
- Design observability, disaster readiness, change management, and service scalability into products.
- Partner with application teams and internal stakeholders to standardize reliability practices.
- Participate in production support and on-call rotations, providing senior-level guidance during critical events.
- Lead incident management and post-incident retrospectives and coach teams in these practices.
- Write design documents, postmortems, and refactor application code.
Requirements
- Experience building automation to reduce operational burden or developing internal SaaS tools.
- Ability to advocate for and introduce SRE principles such as SLOs and SLAs.
- Experience with public cloud or hosted datacenter environments, preferably Azure and AKS.
- Hands-on experience with Linux server stacks, preferably Ubuntu or Debian.
- Knowledge of cloud provisioning platforms, preferably Terraform.
- Exposure to configuration management tools, preferably Chef.
- Experience with containerization or clustering technologies, preferably Docker.
- Familiarity with observability and alerting tools such as Prometheus/Grafana or ELK/EFK.
- Practical experience with CI/CD pipelines and rollout strategies.
- Bachelor’s degree or equivalent experience in Computer Engineering or a related field.
- Proficiency in one or more programming languages such as Java, Python, or Golang.
- Familiarity with scripting languages such as PowerShell, Bash, Python, or Ruby.
- Collaborative communication skills and the ability to influence reliability best practices across teams.
Benefits
- Hybrid work style with Tuesdays and Fridays dedicated to in-office collaboration and Mondays and Fridays described as remote-friendly focus time.
- Flexible work hours.
- Internal career development framework and unlimited access to LinkedIn Learning and Microsoft courses and training.
- Comprehensive health, vision, dental, and life insurance.
- Registered Retirement Savings Plan with company matching up to 5%.
- Annual performance-based bonus.
- Enhanced paid parental leave: 20 weeks for primary leave and 10 weeks for secondary leave.
- Flexible paid time off and company wellness days.
- Access to RethinkCare behavioral health and well-being resources.
- Open-plan workspace with a gaming area, free snacks and drinks, and regular social events.
Categories
DevOpsSite Reliability
