Merative

Principle SRE

Merative
Apply
7 days ago
Remote, United States or Mississauga, CanadaStaff+

Base Salary

$175k - $262k/yr

Responsibilities

  • Define and govern SLOs, SLIs, error budgets, observability standards, and reliability practices across platform and shared services.
  • Own incident detection, response, escalation, post-incident reviews, and corrective actions for complex production issues.
  • Lead capacity planning, performance engineering, failure-domain isolation, disaster recovery design, and resilience improvements.
  • Design and implement infrastructure as code, environment provisioning, deployment automation, automated remediation, and self-service operational tooling.
  • Automate patching, scaling, certificate rotation, backup and restore validation, disaster recovery exercises, and security/compliance evidence collection.
  • Measure and reduce operational toil while establishing measurable toil-reduction targets.
  • Provide technical guidance, mentorship, design-review leadership, and incident-learning support to development, QA, and operations teams.
  • Coordinate with development, product, program management, support, implementation, security, and other cross-functional teams.
  • Plan, track, and deliver reliability and automation initiatives on time and on budget.
  • Represent the platform's reliability posture to customers and auditors and participate in architecture, design, and project documentation.

Requirements

  • Degree in Computer Science, Engineering, or a related field, or equivalent industry-related experience.
  • 10+ years of experience in software engineering, systems engineering, or infrastructure operations with progression into a principal or staff-level technical role.
  • Demonstrated experience informally leading teams, projects, or people to successful outcomes.
  • Experience operating production SaaS at scale against formal availability commitments.
  • Experience designing and implementing operational automation that measurably reduced manual effort or recovery time.
  • Deep expertise in cloud infrastructure, distributed systems, Azure, infrastructure as code, configuration management, containers, orchestration, observability, automation or systems programming, Linux, networking, and identity and authorization.
  • Experience with incident management, post-incident analysis, software architecture, technical project management, and Agile environments.
  • Ability to technically lead and influence experienced professionals without direct authority.
  • Preferred Azure certification such as Azure Solutions Architect Expert or DevOps Engineer Expert.
  • Preferred experience includes regulated environments, medical imaging, DICOM, HL7, service mesh implementation, Kafka, large-scale relational and NoSQL data platforms, cybersecurity, chaos engineering, data engineering, and machine learning or artificial intelligence applied to operations.
  • Strong systems thinking, analytical judgment, communication, collaboration, prioritization, composure during incidents, and automation-first leadership.

Benefits

  • Remote-first work-from-home culture with minimal travel of approximately 5%.
  • Participation in an on-call escalation rotation, including occasional work outside standard business hours.
  • Flexible vacation and paid leave benefits.
  • Health, dental, and vision insurance.
  • 401(k) retirement savings plan.
  • Infertility benefits, tuition reimbursement, life insurance, employee assistance program, and additional benefits.

Tech Stack

Categories

Site Reliability
Merative

About Merative

1,001-5,000 employees
Contact me