IFS

Principal Site Reliability Engineer – Performance (A&D, Ultra HA/Exadata)

IFS
Apply
12 days ago
Ottawa, CanadaStaff+

Responsibilities

  • Define technical direction, SRE operating models, reliability architecture, SLOs, error budgets, incident management, capacity planning, and disaster recovery strategies for the A&D vertical.
  • Lead enterprise-scale monitoring, observability, automated remediation, infrastructure-as-code, CI/CD, Kubernetes, multi-cloud orchestration, and cloud-native architecture initiatives.
  • Own Ultra HA performance and availability on Oracle Exadata, including database health monitoring, workload and wait-event analysis, application performance management, tuning, capacity planning, and SLA protection.
  • Design and maintain performance observability using Elastic, Grafana, OpenTelemetry, APM, dashboards, baselines, and alert thresholds.
  • Lead performance investigations, load testing, pre-release validation, customer reporting, and service reviews for demanding enterprise workloads.
  • Lead complex and critical incidents, establish incident command and escalation procedures, conduct blameless post-incident reviews, and eliminate recurring causes.
  • Drive automation-first practices, self-healing systems, predictive alerting, runbook automation, GitOps, and reductions in MTTD and MTTR.
  • Create architecture documentation, disaster recovery playbooks, operational runbooks, knowledge bases, and training programs.
  • Partner with R&D, Product, Operations, Customer Success, and Professional Services to translate reliability requirements into technical strategies.
  • Manage and mentor 3–5 SRE engineers while establishing technical standards, career progression, and sustainable on-call practices.

Requirements

  • A university degree or equivalent professional qualification in Software Engineering, Computer Science, Information Technology, or a related discipline.
  • At least 10 years of progressive experience in cloud computing services, enterprise IT delivery, or site reliability engineering, including 7–8 years in hands-on SRE, DevOps, or cloud operations roles.
  • Experience leading SRE teams, establishing SRE practices, driving organizational change, and implementing large-scale automation platforms.
  • Expertise in cloud infrastructure and operations across Azure, AWS, and GCP, plus production-scale Kubernetes and Docker operations.
  • Hands-on production experience with Oracle Exadata and expert Oracle database administration, monitoring, diagnostics, performance tuning, high availability, backup, recovery, and disaster recovery.
  • Experience with Oracle tools and technologies including AWR, ASH, ADDM, Statspack, SQL Trace/TKPROF, Oracle Enterprise Manager, RAC, Data Guard, Active Data Guard, ASM, RMAN, Smart Scan, storage indexes, IORM, flash cache, and cell-level diagnostics.
  • Experience with performance engineering, load testing, capacity planning, latency and throughput analysis, application optimization, and mission-critical high-availability systems.
  • Experience administering and troubleshooting WildFly, WebLogic, Nginx, Linux systems, networks, application servers, JVMs, and multi-tier applications.
  • Advanced scripting skills in Bash, PowerShell, Python, or Go and infrastructure-as-code experience with Terraform, Ansible, or CloudFormation.
  • Experience designing CI/CD pipelines, GitOps practices, enterprise monitoring, Elasticsearch or the Elastic Stack, Grafana, Prometheus, OpenTelemetry, and APM tooling such as Dynatrace, AppDynamics, New Relic, or Elastic APM.
  • Expertise debugging complex applications, analyzing network packets and protocols, and correlating database, infrastructure, and application telemetry.
  • Service desk tooling experience with ServiceNow, Jira Service Desk, or an equivalent; ITIL, ISO 20000, or equivalent service-delivery certification or demonstrated practical mastery.
  • Preferred qualifications include cloud security, Zero Trust, IAM, encryption, FinOps, multi-region disaster recovery, Kafka, RabbitMQ, API management, machine learning or AI for observability, and certifications such as AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA, OCP, Exadata, Elastic, or Grafana certifications.
  • Strong strategic thinking, executive communication, technical leadership, mentoring, organizational influence, decision-making under pressure, collaboration, and English communication skills.

Benefits

  • Global, diverse, and inclusive work environment with opportunities for worldwide impact.
  • Opportunity to work on AI-driven enterprise software, mission-critical Aerospace & Defense systems, and large-scale reliability transformation.
  • Role includes leadership and mentoring responsibilities, technical training, knowledge sharing, and participation in SRE communities of practice.

Tech Stack

AnsibleApache KafkaAssemblyAWSAzureBashDatadogDockerElasticsearchGoGoogle Cloud PlatformGrafanaKibanaKubernetesLogstashOracle DatabasePowerShellPrometheusPythonRabbitMQSplunkTerraform

Categories

DevOpsSite Reliability
IFS

About IFS

5,001-10,000 employees
Contact me