
Lead, Site Reliability Engineering (Application Support)
Ontario Municipal Employees Retirement System15 days ago
Toronto, CanadaStaff+
Responsibilities
- Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.
- Respond to incidents, lead triage, and collaborate with SRE, platform, network, security, and application teams to restore services.
- Support deployments, releases, change management, and GitHub Actions CI/CD workflows.
- Resolve access issues involving Azure AD groups, SSO, application permissions, and firewall rules.
- Configure and support Azure platform components, networking, certificates, shared cloud services, and observability tooling.
- Onboard applications and teams to the DEV platform, including setup, access, deployment readiness, monitoring, and operational handover.
- Develop runbooks, support procedures, knowledge articles, and operational documentation.
- Contribute to automation and reliability improvements while providing technical guidance on SRE and platform support practices.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.
- Strong hands-on experience with Microsoft Azure services including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management, and Azure Functions.
- Experience supporting production environments, incident response, troubleshooting, problem management, and operational support processes.
- Experience with container technologies, cloud-native application architectures, CI/CD pipelines, and deployment automation using GitHub Actions or similar platforms.
- Knowledge of identity, networking, and access management concepts including SSO, OAuth, application registrations, and security groups.
- Experience with Datadog, Azure Monitor, Log Analytics, or comparable observability and monitoring platforms.
- Understanding of DNS, certificates, firewalls, private endpoints, and network security controls.
- Experience with scripting and automation using PowerShell, Bash, Azure CLI, Python, or similar tools.
- Strong knowledge of operating systems and cloud infrastructure concepts, with effective cross-functional communication and problem-solving skills.
- Preferred qualifications include Kubernetes or container app experience, enterprise developer platform support, SRE principles such as SLOs, SLIs, and error budgets, enterprise API integrations, Azure networking and security practices, AI/LLM or agent-platform exposure, relevant certifications, and post-secondary education in a related discipline.
Benefits
- Flexible hybrid work arrangement requiring teams to work in the office 4 days per week.
- Eligibility for annual incentive awards under short-term and, if applicable, long-term incentive plans.
- Participation in group benefits and retirement plans.
- Inclusive, barrier-free recruitment and selection process with employee resource groups, recognition programs, and wellness and development initiatives.
Tech Stack
Categories
DevOpsSite Reliability