3 months ago
Remote, Canada or Vancouver, CanadaMid Level
Responsibilities
- Ensure the reliability of critical products and services by meeting or exceeding SRE objectives.
- Instantiate and maintain production infrastructure using Infrastructure as Code and configuration management tools.
- Build and maintain centralized logging, monitoring, time-series databases, and alerting systems.
- Automate service deployments, administration, and monitoring using CI/CD practices.
- Collaborate with engineering and information security teams to improve service operability, documentation, processes, and security.
- Participate in the team on-call rotation and support incident management, troubleshooting, postmortems, and root cause analysis.
Requirements
- At least 3 years of software and/or operational experience building and maintaining internet-facing production environments.
- Bachelor’s degree in information systems, computer science, technology, or a related field is strongly preferred, or 2+ years of relevant and/or equivalent experience in lieu of a degree.
- Strong experience administering Linux/Unix systems.
- Experience with source control tools, preferably Git.
- Experience with configuration management and Infrastructure as Code tools, preferably Ansible, Puppet, and Terraform.
- Understanding of container technology, preferably Docker and Kubernetes.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, Nagios, or similar systems.
- Experience with non-cloud infrastructure and large-scale 24/7 production environments.
- Strong scripting abilities in Bash and Python.
- Experience with incident management, troubleshooting, root cause analysis, postmortems, incident response plans, and incident resolution improvement.
- Experience maintaining build systems such as Jenkins, DroneCI, or similar tools.
- Experience across the full software lifecycle, including systems architecture, systems design, implementation, maintenance, and operations.
- Programming experience using HTTP Service APIs.
- Virtualization experience with VMWare, Proxmox, or Oracle Linux Virtualization Manager.
- Network administration experience, distributed data processing, databases, large-scale file systems, security and testing frameworks, or regulated-industry environments are pluses.
Benefits
- Full-time remote position from British Columbia on the Pacific time zone.
- Inclusive workplace with equal-opportunity employment and career-development support.
Categories
DevOpsSite Reliability
