
Site Reliability Engineer (SRE) I
Thomson Reuters24 hours ago
Base Salary
$71k - $131k/yr
Responsibilities
- Maintain SRE operational tooling, including dashboards, alerts, runbooks, service documentation, telemetry baselines, deployment visibility, and dependency information.
- Use logs, metrics, traces, dashboards, and alerts to investigate service-health issues and support incident response.
- Execute approved runbooks and mitigation procedures within established escalation and change-management processes.
- Document incident findings, handoffs, operational reviews, and open questions clearly.
- Contribute code, scripts, infrastructure configuration, dashboards, alerts, automation, and documentation that improve reliability and reduce operational toil.
- Improve monitoring, observability, deployment visibility, service ownership records, and operational workflows.
- Assist with root-cause analysis, post-incident reviews, corrective actions, and service-health initiatives involving SLOs and error budgets.
- Use and validate AI-enabled coding, documentation, investigation, and incident-management tools.
- Partner with product engineering and platform teams on reliability, observability, capacity, alert quality, and failure-mode improvements.
Requirements
- At least 3 years of experience in SRE, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, software engineering, or a related technical field.
- Working knowledge of at least two areas including cloud infrastructure, distributed systems, observability and monitoring, networking, databases, CI/CD, containers, or infrastructure automation.
- Experience with production telemetry, production troubleshooting or incident response, and business-critical applications or services.
- Experience writing, maintaining, or improving operational documentation, runbooks, knowledge articles, or support procedures.
- Experience scripting or programming in Python, Bash, PowerShell, JavaScript, Java, Go, or a comparable language.
- Ability to contribute to code, infrastructure configuration, dashboards, monitoring rules, alerts, automation, or documentation.
- Clear communication of technical findings, operational risks, and next steps.
- Familiarity with AI-enabled engineering and operational-analysis tools and the ability to review and validate their outputs.
- Preferred experience with cloud platforms, Kubernetes, containers, CI/CD tooling, infrastructure-as-code, configuration management, observability platforms, SLOs, SLIs, error budgets, incident management, 24/7 production environments, on-call rotations, and blameless post-incident reviews.
- Relevant certifications in cloud infrastructure, Kubernetes, DevOps, SRE, observability, or incident management are beneficial but not required.
Benefits
- Flexible hybrid work model for office-based roles.
- Work-from-anywhere flexibility for up to 8 weeks per year and other work-life balance policies.
- Career development, continuous learning, Grow My Way programming, and skills development support.
- Health, dental, vision, disability, and life insurance; 401(k) with company match; vacation, sick and safe paid time off; paid holidays; parental and sabbatical leave.
- Two company-wide mental health days, Headspace access, fitness reimbursement, Employee Assistance Program, tuition reimbursement, commuter benefits, and employee incentive programs.
- Two paid volunteer days annually and opportunities for pro-bono consulting and ESG initiatives.
Tech Stack
Categories
Site Reliability
About Thomson Reuters
Thomson Reuters builds research platforms, workflow software, and data services for legal, tax and accounting, compliance, and government professionals, and operates the Reuters global news service. Flagship products include Westlaw for legal research and ONESOURCE and Checkpoint for tax and accounting, sold primarily via subscriptions and enterprise licenses. A public company headquartered in Toronto, it was formed in 2008 by combining Thomson Corporation and Reuters Group and is listed on the TSX and Nasdaq as TRI.