
Systems Engineer, Senior - Observability
Mass General Brigham2 months ago
Somerville, MA, USASenior
Base Salary
$94k - $137k/yr
Responsibilities
- Deploy, configure, optimize, and maintain Dynatrace and Cisco ThousandEyes observability platforms.
- Build dashboards, alerts, management zones, tagging rules, synthetic tests, and network-path tests.
- Develop Terraform configuration-as-code using Azure DevOps repositories and pipelines, with a planned transition to GitHub Enterprise.
- Create and maintain custom Dynatrace extensions and implement observability for Azure DevOps frameworks and CI/CD pipelines.
- Onboard applications and infrastructure through instrumentation, monitoring configuration, and data validation.
- Automate alerting and remediation workflows to improve mean time to resolution and service uptime.
- Establish observability standards with application, cloud, network, and security teams.
- Create documentation, operational guidance, and best practices; mentor team members and support technical outcomes.
- Participate in an on-call rotation for observability platform health and incident response.
Requirements
- Bachelor’s degree in Computer Science or a related field, or relevant experience in lieu of a degree.
- 5–7 years of experience as a systems engineer or in a related technical engineering role.
- Hands-on Dynatrace experience covering application performance monitoring, infrastructure monitoring, real user monitoring, and Dynatrace Query Language.
- Experience with Cisco ThousandEyes for synthetic testing, path visualization, and internet or wide area network performance monitoring.
- Experience using Terraform and the Dynatrace Terraform provider for configuration-as-code.
- Experience with Azure DevOps source control and pipelines; familiarity with GitHub Enterprise is helpful.
- Strong Python and JavaScript programming skills for automation, custom telemetry, and tooling.
- Experience developing custom Dynatrace extensions and implementing observability for Azure DevOps and CI/CD pipelines.
- Foundational knowledge of applications, servers, storage, and networks, plus strong troubleshooting skills across metrics, logs, and traces.
- Clear communication, cross-team collaboration, vendor coordination, and mentoring abilities.
- Preferred: PowerShell, Splunk, Microsoft SCOM, SiteScope, Nagios, Monaco, Dynatrace Grail, Dynatrace or Azure certifications, and familiarity with Site Reliability Engineering practices.
Benefits
- Full-time Monday–Friday schedule during Eastern business hours with 40 scheduled weekly hours.
- Hybrid work model with on-site work at local Mass General Brigham sites weekly or monthly as business needs require.
- Required on-call rotation, typically one week at a time, supporting platform health and incident response.
- Periodic in-person stakeholder and team meetings; remote work requires a stable, secure, compliant workstation and Microsoft Teams participation using MGB-provided equipment.
- Comprehensive benefits, career advancement opportunities, differentials, premiums, bonuses as applicable, and recognition programs.
Tech Stack
Categories
DevOpsSite Reliability