7 days ago
Remote, United States or Mississauga, CanadaStaff+
Base Salary
$175k - $262k/yr
Responsibilities
- Define and govern SLOs, SLIs, error budgets, observability standards, and reliability practices across platform and shared services.
- Own incident detection, response, escalation, post-incident reviews, and corrective actions for complex production issues.
- Lead capacity planning, performance engineering, failure-domain isolation, disaster recovery design, and resilience improvements.
- Design and implement infrastructure as code, environment provisioning, deployment automation, automated remediation, and self-service operational tooling.
- Automate patching, scaling, certificate rotation, backup and restore validation, disaster recovery exercises, and security/compliance evidence collection.
- Measure and reduce operational toil while establishing measurable toil-reduction targets.
- Provide technical guidance, mentorship, design-review leadership, and incident-learning support to development, QA, and operations teams.
- Coordinate with development, product, program management, support, implementation, security, and other cross-functional teams.
- Plan, track, and deliver reliability and automation initiatives on time and on budget.
- Represent the platform's reliability posture to customers and auditors and participate in architecture, design, and project documentation.
Requirements
- Degree in Computer Science, Engineering, or a related field, or equivalent industry-related experience.
- 10+ years of experience in software engineering, systems engineering, or infrastructure operations with progression into a principal or staff-level technical role.
- Demonstrated experience informally leading teams, projects, or people to successful outcomes.
- Experience operating production SaaS at scale against formal availability commitments.
- Experience designing and implementing operational automation that measurably reduced manual effort or recovery time.
- Deep expertise in cloud infrastructure, distributed systems, Azure, infrastructure as code, configuration management, containers, orchestration, observability, automation or systems programming, Linux, networking, and identity and authorization.
- Experience with incident management, post-incident analysis, software architecture, technical project management, and Agile environments.
- Ability to technically lead and influence experienced professionals without direct authority.
- Preferred Azure certification such as Azure Solutions Architect Expert or DevOps Engineer Expert.
- Preferred experience includes regulated environments, medical imaging, DICOM, HL7, service mesh implementation, Kafka, large-scale relational and NoSQL data platforms, cybersecurity, chaos engineering, data engineering, and machine learning or artificial intelligence applied to operations.
- Strong systems thinking, analytical judgment, communication, collaboration, prioritization, composure during incidents, and automation-first leadership.
Benefits
- Remote-first work-from-home culture with minimal travel of approximately 5%.
- Participation in an on-call escalation rotation, including occasional work outside standard business hours.
- Flexible vacation and paid leave benefits.
- Health, dental, and vision insurance.
- 401(k) retirement savings plan.
- Infertility benefits, tuition reimbursement, life insurance, employee assistance program, and additional benefits.
Tech Stack
Categories
Site Reliability
