
RPE - System Reliability Engineering Specialist (Hybrid)
Marks & Spencer Group plc1 hour ago
Montréal, CanadaMid Level
Responsibilities
- Support business-critical applications and resolve ambiguous operational and technical problems.
- Manage major incidents, perform root cause analysis, and implement corrective actions to improve reliability.
- Develop and maintain automation, scripts, operational tooling, documentation, and runbooks.
- Perform platform upgrades, patching, archival, cleanup, and ongoing maintenance.
- Evaluate, deploy, and operate third-party software, AI-enabled platforms, and emerging technologies.
- Collaborate with development, QA, product management, business stakeholders, and global teams on application support and continuous improvement.
Requirements
- At least 2 years of relevant experience.
- Strong knowledge of Linux/Unix operating systems and production support environments.
- Strong scripting and automation skills using Python, Bash, Shell scripting, or similar technologies.
- Understanding of networking concepts, protocols, APIs, REST services, system integrations, and troubleshooting methodologies.
- Experience with monitoring, observability, and reliability engineering tools such as OpenTelemetry, Grafana, and Loki.
- Experience with Visual Studio Code and Git.
- Familiarity with Docker, Kubernetes, cloud technologies, and automation frameworks.
- Basic understanding of generative AI concepts including Large Language Models, AI agents, prompt engineering, and Retrieval-Augmented Generation.
- Familiarity with AI-assisted development tools such as GitHub Copilot or Microsoft Copilot.
- Knowledge of French and English is required.
Benefits
- Hybrid work environment combining remote work with attendance in the Montreal, Quebec office.
- Comprehensive employee benefits and perks for employees and their families.
- Equal opportunity workplace with opportunities for career mobility across the business.
Tech Stack
Categories
Site Reliability