
Staff Site Reliability Engineer
Johnson Controls International plc1 hour ago
Richmond Hill, CanadaStaff+
Responsibilities
- Own escalated production issues across the OpenBlue Data Platform, enterprise SaaS products, and Airwall.
- Debug cross-layer failures spanning application code, data pipelines, cloud infrastructure, and network paths.
- Lead root cause analysis and drive interim and permanent corrective actions through closure.
- Improve instrumentation, monitors, dashboards, runbooks, alert quality, and production detection using Datadog and Grafana.
- Participate in the on-call rotation and lead technical response during major incidents.
- Represent Engineering in customer conversations and post-incident reviews for high-severity incidents.
- Plan and execute infrastructure upgrades, migrations, and platform changes across Azure and AWS.
- Own Terraform infrastructure as code, including module design, state management, and drift remediation.
- Operate and improve Kubernetes workloads, including capacity, autoscaling, resource limits, and deployment reliability.
- Improve release and change safety, reduce manual toil through automation, and establish operational standards for observability, change management, and production readiness.
- Harden environments serving government and data-residency-sensitive customers in partnership with security and compliance teams.
Requirements
- Must reside in Canada and be legally authorized to work in Canada without sponsorship.
- At least 7 years of experience in site reliability engineering, production engineering, L3 application support, or infrastructure engineering.
- Demonstrated ability to take complex production problems from investigation through permanent remediation.
- Strong hands-on Terraform experience with production ownership of infrastructure as code.
- Production experience with Azure, AWS, and Kubernetes at scale.
- Practical experience with Datadog and Grafana, including instrumentation, dashboards, monitor design, and alert quality.
- Experience using AI-native development tools appropriately, including Claude, Copilot, Codex, or Cursor.
- Willingness to participate in on-call rotations and respond to incidents outside business hours.
- Working proficiency in Java and C# sufficient to read, diagnose, and correct application code.
- Experience supporting government or public-sector customers, including data residency, data sovereignty, or Protected B requirements, is preferred.
- Eligibility to obtain Government of Canada security screening at Reliability Status or higher is preferred.
- Experience operating data platforms with streaming and batch pipelines, data-quality monitoring, and latency service-level objectives is preferred.
- Familiarity with operational technology networking and zero-trust network architecture is preferred.
- Prior L2 or L3 support organization experience with formal service-level agreements and escalation structures is preferred.
- Exposure to building automation, HVAC, or connected-building technology is preferred.
Benefits
- The position is based in Canada and works directly with Canadian-resident systems and data.
- The role is onsite, as indicated by the #LI-ONSITE designation.
- The position includes participation in an on-call rotation and response to incidents outside business hours.
- Johnson Controls provides reasonable accommodation throughout the recruitment and selection process for applicants, candidates, and employees with disabilities.
Tech Stack
Categories
Site Reliability