RBC

Site Reliability Engineer (SRE), Cloud Operations

RBC
Apply
1 day ago
Toronto, CanadaSenior

Responsibilities

  • Support scalable, secure, and highly available architectures across private and public cloud platforms.
  • Write code and scripts to automate infrastructure workflows and reduce operational toil across Data Center, Branch, and NOC operations.
  • Extend Ansible-based self-healing automation for routine remediation tasks such as CPU remediation.
  • Participate in and lead design reviews for platform features, infrastructure changes, and operational integrations.
  • Provide technical feedback, contribute code changes to shared repositories, and establish operational data standards and pipelines.
  • Drive automation, CI/CD, and Infrastructure as Code practices using Ansible and Terraform.
  • Use proactive alerting and anomaly detection to reduce reliability risks involving durability, availability, performance, and correctness.
  • Participate in on-call rotations, incident management, troubleshooting, and incident triage using observability and paging tools.

Requirements

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong knowledge of Kubernetes and OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting for infrastructure automation and platform development.
  • Hands-on experience with monitoring and observability stacks such as Prometheus, Grafana, and ELK.
  • Experience with incident management, on-call rotations, and post-incident review practices.
  • Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations, including anomaly detection, intelligent alerting, and predictive capacity planning.
  • Preferred experience with AWS, Azure, or GCP in hybrid or multi-cloud environments.
  • Preferred experience with GPU and compute infrastructure for ML inference workloads.

Benefits

  • Comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable.
  • Development support through coaching and management opportunities.
  • Flexible work/life balance options.
  • Opportunities for challenging work and progressively greater accountabilities.
  • Access to varied career opportunities across the business.
  • Full-time regular employment based in Toronto with 37.5 work hours per week.

Categories

DevOpsSite Reliability
RBC

About RBC

10,000+ employees

Royal Bank of Canada (RBC) provides personal and commercial banking, payments, mortgages, investing, wealth management, insurance, and capital markets services to consumers, businesses, and institutions. The public company, headquartered in Toronto and founded in 1864, serves more than 17 million clients across Canada, the U.S., and other global markets. Revenue comes from interest income, fees, advisory, and trading across its diversified financial services businesses, and it is Canada’s largest bank by market capitalization.

Contact me