Base Salary
$80k - $140k/yr
Responsibilities
- Build and enhance intelligent monitoring, alerting, reliability testing, anomaly detection, and automated remediation capabilities.
- Deploy metrics, logs, traces, dashboards, and actionable alerting across supported applications.
- Design ML-based anomaly detection and self-healing solutions with appropriate governance controls.
- Standardize telemetry and instrumentation across distributed platforms.
- Automate operational workflows using Ansible, GitHub Actions, Bash, Python, PowerShell, and custom tooling.
- Define and improve SLIs, SLOs, error budgets, alerting strategies, service health metrics, and automation-first remediation runbooks.
- Partner with development teams to ensure applications meet reliability and performance standards before and after deployment.
- Lead incident and problem management, participate in an on-call rotation, troubleshoot production issues, and drive root cause analysis and corrective actions.
- Identify opportunities to simplify, automate, and modernize operations using engineering and AI-driven approaches.
Requirements
- 5+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Systems Engineering roles with strong operational depth.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong experience with infrastructure automation and configuration management, particularly Ansible.
- Strong scripting and automation skills in Bash, Python, PowerShell, or similar languages.
- Hands-on experience with reliability and observability tools such as Elasticsearch, Dynatrace, GitHub, Kubernetes, OpenShift, Kafka, PagerDuty, Moogsoft, or related platforms.
- Strong understanding of production operations, incident management, root cause analysis, and reliability engineering practices.
- Experience defining and operating SLIs, SLOs, alerting strategies, and service health metrics.
- Knowledge of cloud-native and distributed-systems concepts, including resiliency, scalability, fault isolation, and performance tuning.
- Understanding of AIOps, AI/ML concepts, or intelligent automation applied to observability and operations.
- Ability to work across teams, influence engineering practices, and communicate with technical and non-technical stakeholders.
- Experience in financial services, wealth management, banking, insurance, or other highly regulated environments is preferred.
- Experience with OpenTelemetry and telemetry standardization across distributed systems is preferred.
- Experience with Prometheus, Grafana, Splunk, Catchpoint, Azure Automation, or similar SRE and observability platforms is preferred.
- Experience with CI/CD and developer platform tools such as Jenkins, Artifactory, and Vault is preferred.
- Familiarity with Docker, Kubernetes-based deployments, anomaly detection, predictive alerting, self-healing automation, AI governance, model validation, and operational controls is preferred.
Benefits
- Comprehensive Total Rewards Program including bonuses and flexible benefits.
- Competitive compensation, commissions, and stock where applicable.
- Discretionary variable compensation may provide additional total compensation based on business and individual performance.
- Development support through coaching and management opportunities.
- Opportunity to make a difference and lasting impact.
- Collaborative, progressive, high-performing team environment.
- World-class training program in financial services.
- Full-time, regular salaried position working 40 hours per week in Minneapolis.
Tech Stack
Categories
About RBC
Royal Bank of Canada is a global financial institution with a purpose-driven, principles-led approach to delivering leading performance. Our success comes from the 94,000+ employees who leverage their imaginations and insights to bring our vision, values and strategy to life so we can help our clients thrive and communities prosper. As Canada's biggest bank and one of the largest in the world, based on market capitalization, we have a diversified business model with a focus on innovation and providing exceptional experiences to our more than 17 million clients in Canada, the U.S. and 27 other countries. Learn more at rbc.com. We are proud to support a broad range of community initiatives through donations, community investments and employee volunteer activities. See how at www.rbc.com/community-social-impact. http://rbc.com/legalstuff. La Banque Royale du Canada est une institution financière mondiale définie par sa raison d'être, guidée par des principes et orientée vers l'excellence en matière de rendement. Notre succès est attribuable aux quelque 94 000+ employés qui mettent à profit leur créativité et leur savoir faire pour concrétiser notre vision, nos valeurs et notre stratégie afin que nous puissions contribuer à la prospérité de nos clients et au dynamisme des collectivités. Selon la capitalisation boursière, nous sommes la plus importante banque du Canada et l'une des plus grandes banques du monde. Nous avons adopté un modèle d'affaires diversifié axé sur l'innovation et l'offre d'expériences exceptionnelles à nos plus de 17 millions de clients au Canada, aux États Unis et dans 27 autres pays. Pour en savoir plus, visitez le site rbc.com/francais Nous sommes fiers d'appuyer une grande diversité d'initiatives communautaires par des dons, des investissements dans la collectivité et le travail bénévole de nos employés. Pour de plus amples renseignements, visitez le site www.rbc.com/collectivite-impact-social. https://www.rbc.com/conditions-dutilisation/
