
Senior Observability Engineer
CI Financial Corp.5 days ago
Toronto, CanadaSenior
Base Salary
$85k - $125k/yr
Responsibilities
- Design, deploy, optimize, and govern enterprise Dynatrace observability platforms across AWS and hybrid environments.
- Configure APM instrumentation, distributed tracing, dashboards, alerting, synthetic monitoring, log monitoring, and telemetry standards.
- Support complex incident triage and post-incident reviews using Dynatrace diagnostics and observability evidence.
- Define and maintain SLIs, SLOs, error budgets, monitoring coverage, and reliability improvements.
- Automate monitoring configurations and platform operations using Terraform or Dynatrace Monaco, Python, Bash, PowerShell, REST APIs, and webhooks.
- Lead legacy monitoring migrations, close monitoring gaps, and integrate observability into application and infrastructure delivery workflows.
- Manage telemetry quality, retention, platform usage, alert design, reporting, and Dynatrace consumption costs.
- Build reliability and service-health dashboards for engineering and executive audiences.
- Mentor engineers and partner teams on instrumentation, onboarding, dashboarding, and observability best practices.
Requirements
- 5–10 years of experience in observability, monitoring engineering, SRE, APM, DevOps, or infrastructure engineering, including several years of hands-on Dynatrace administration and architecture.
- Deep hands-on expertise with Dynatrace full-stack monitoring, Davis AI, Smartscape, RUM, synthetic monitoring, distributed tracing, Grail, DQL, management zones, Workflows/AutomationEngine, and access governance.
- Strong experience with AWS services and architecture, especially CloudWatch, EKS, ECS, Lambda, EC2, RDS, API Gateway, and networking.
- Advanced knowledge of Kubernetes and cloud-native observability for microservices and distributed systems.
- Proficiency with observability-as-code using Terraform and/or Monaco and scripting with Python, Bash, or PowerShell.
- Understanding of distributed application architecture, networking, telemetry pipelines, performance engineering, and incident management.
- Dynatrace Associate or Professional certification is required; higher-level certification is preferred.
- Experience with Nagios, SolarWinds, Prometheus, Grafana, Splunk, ELK, OpenTelemetry, ServiceNow, or other ITSM/event-management platforms is preferred.
- Background in SRE practices such as reliability reviews, error-budget management, and incident-reduction programs is preferred.
Benefits
- In-office work environment with a modern headquarters within walking distance of Union Station.
- Health insurance, enhanced group benefits, wellness programs, life and disability insurance, and retirement savings plans.
- Paid leave, paid holidays, vacation time, parental-leave top-up, and paid time off for volunteering.
- Training reimbursement and paid professional designations.
- Employee Savings Plan and corporate discount program.