3 months ago
Ottawa, CanadaMid Level
Responsibilities
- Ensure the reliability of critical products and services by meeting or exceeding SRE objectives.
- Instantiate and maintain production infrastructure using Infrastructure as Code and Configuration Management tools.
- Build and maintain centralized logging, time-series monitoring, and alerting systems.
- Automate service deployments, administration, and monitoring using CI/CD practices.
- Collaborate with engineering and information security teams to improve, document, and establish service operability and security processes.
- Participate in the team on-call rotation and handle incident management, troubleshooting, root cause analysis, postmortems, and incident response improvements.
- Run and maintain build systems and support the full software lifecycle from architecture and design through implementation, maintenance, and operations.
Requirements
- At least 3 years of software and/or operational experience building and maintaining internet-facing production environments.
- Bachelor’s degree in information systems, computer science, technology, or a related field is strongly preferred; 2+ years of relevant or equivalent experience may substitute for the degree.
- Strong experience with Linux/Unix systems administration.
- Experience with source control tools, preferably Git.
- Experience with Configuration Management and Infrastructure as Code tools, preferably Ansible, Puppet, and Terraform.
- Understanding of container technology, preferably Docker and Kubernetes.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, or Nagios.
- Experience with non-cloud infrastructure and large-scale 24/7 production environments.
- Strong scripting abilities in Bash and Python.
- Experience with incident management, troubleshooting, root cause analysis, postmortems, incident response plans, and incident resolution procedures.
- Experience maintaining build systems such as Jenkins or DroneCI.
- Programming experience using HTTP Service APIs and experience with virtualization technologies such as VMWare, Proxmox, or Oracle Linux Virtualization Manager.
- Network administration experience, exposure to security and testing frameworks, and experience with distributed data processing, databases, large-scale file systems, or regulated industries are additional qualifications.
Benefits
- Full-time hybrid position requiring attendance at the Ottawa office at least 3–4 days per week.
- Target compensation package of CAD $80,000–$100,000, subject to internal equity and years of experience.
- Inclusive workplace with career growth and promotion opportunities.
Categories
DevOpsSite Reliability
