
Incident Commander
Penn Interactive12 days ago
Remote, CanadaSenior
Responsibilities
- Lead real-time incident management across engineering, application, customer support, SRE, infrastructure, and other cross-functional teams.
- Classify, document, investigate, escalate, diagnose, recover, and perform root-cause analysis for P1, P2, P3, and P4 incidents.
- Develop and maintain incident-management practices, frameworks, process flows, templates, process guides, tools, and practical capabilities.
- Collaborate with SRE and infrastructure teams to identify requirements and gaps that contribute to downtime or operational blind spots.
- Drive continuous service improvements and improvements to service delivery and release processes based on disruption reports.
- Maintain root-cause analysis documentation.
- Lead timely SRE communications to stakeholders through email, Slack, and Microsoft Teams.
- Promote quality and alignment in JIRA release-ticket management and incident-management communications in support of service-level agreements.
Requirements
- Experience in a similar incident-management role.
- Experience and understanding of containerization, preferably Docker and Kubernetes.
- Understanding of configuration management and infrastructure-as-code tools such as Terraform, Ansible, and Helm.
- Experience with a programming language and comfort working in Linux environments.
- Experience with AWS, Google Cloud Platform, and on-premises environments.
- Experience working in a 24x7 on-call environment and the ability to manage multiple projects independently.
- Degree in computer science, engineering, or a similar field, or similar experience.
- Preferred experience with PostgreSQL, MySQL, Elasticsearch, Kafka, Redis, Terragrunt, Prometheus, Python, and Talos Linux.
Benefits
- Remote role.
- Competitive compensation package with a CAD $90,000–$135,000 base salary range.
- Education and conference reimbursements.
- Parental leave top-up.
- Career progression opportunities and mentoring opportunities.
- Fun, relaxed work environment.
Tech Stack
AnsibleApache KafkaAWSDockerElasticsearchGoogle Cloud PlatformHelmKubernetesLinuxMySQLPostgreSQLPrometheusPythonRedisTerraform
Categories
Site Reliability