
Senior Infrastructure SRE
PointClickCare8 hours ago
Mississauga, CanadaSenior
Responsibilities
- Design and implement highly available infrastructure for compute, storage, identity, messaging, and shared services.
- Build and maintain Infrastructure as Code using Terraform or Pulumi and establish reusable standards and modules.
- Automate operational workflows, including auto-remediation, self-healing systems, capacity planning, and toil reduction.
- Define and manage SLIs, SLOs, error budgets, reliability targets, and toil metrics.
- Participate in on-call rotations, lead complex infrastructure incident response, conduct blameless post-mortems, and drive systemic fixes.
- Develop observability strategies covering metrics, logs, distributed tracing, and alerting.
- Apply AI-assisted tools to log analysis, alert triage, runbook drafting, post-mortem drafting, and automation scaffolding.
- Own infrastructure-domain reliability end to end and lead multi-team initiatives through adoption.
- Mentor intermediate SREs, review infrastructure changes, and establish operational best practices.
Requirements
- 5+ years of hands-on experience operating and designing cloud infrastructure.
- Expert-level knowledge of Azure or AWS and working proficiency in at least one additional platform among Azure, AWS, or GCP.
- Experience designing and supporting production infrastructure spanning multiple cloud platforms.
- 3+ years of production Infrastructure as Code experience with Terraform, Pulumi, or CloudFormation.
- Ability to design scalable IaC modules, manage multi-cloud state and provider differences, and enforce GitOps workflows.
- Strong proficiency in Python, Go, or Bash for tested, maintainable production automation.
- Production experience applying SRE principles, defining SLIs and SLOs, managing error budgets, and improving reliability.
- Expertise running Kubernetes and containerized workloads on managed Kubernetes platforms such as AKS or EKS, plus VM-based compute.
- Practical production experience operating a service mesh, preferably Istio, or Linkerd or an equivalent.
- Strong proficiency with SAML, OAuth/OIDC, LDAP, cloud IAM, and enterprise SSO; familiarity with PingFederate, Entra ID, Okta, or ADFS.
- Strong proficiency with object, block, and file storage and Kubernetes persistent volumes.
- Experience operating infrastructure in regulated environments and familiarity with HIPAA, SOC 2, PCI, FedRAMP, audit evidence, access controls, encryption, and data residency.
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related technical field, or equivalent practical experience.
- Preferred qualifications include 2+ years in healthcare technology or regulated SaaS, cloud certifications, multi-cluster Kubernetes experience, CI/CD and deployment automation experience, AI-assisted operations experience, and open-source contributions.
Benefits
- Benefits begin on Day 1 and include retirement plan matching, flexible paid time off, wellness programs, parental and caregiver leaves, fertility and adoption support, continuous development support, employee assistance, inclusion communities, and employee recognition.
- The role is hybrid, with an expectation to reside within commuting distance of the specified office and attend regular team events; the listing references the Mississauga and/or Salt Lake City offices.
- The company provides accommodations for candidates participating in the selection process.
Tech Stack
Categories
DevOpsSite Reliability
About PointClickCare
PointClickCare builds a cloud-based EHR and care coordination platform for long-term and post-acute care providers, hospitals, and health plans. Its subscription SaaS covers clinical documentation, eMAR and point-of-care workflows, interoperability, and revenue cycle, and supports a marketplace of integrated partners. Founded in 2000 and headquartered in Toronto, it is privately held and serves over 27,000 providers across North America.