Mass General Brigham

Systems Engineer, Senior - Observability

Mass General Brigham
Apply
2 months ago
Somerville, MA, USASenior

Base Salary

$94k - $137k/yr

Responsibilities

  • Deploy, configure, optimize, and maintain Dynatrace and Cisco ThousandEyes observability platforms.
  • Build dashboards, alerts, management zones, tagging rules, synthetic tests, and network-path tests.
  • Develop Terraform configuration-as-code using Azure DevOps repositories and pipelines, with a planned transition to GitHub Enterprise.
  • Create and maintain custom Dynatrace extensions and implement observability for Azure DevOps frameworks and CI/CD pipelines.
  • Onboard applications and infrastructure through instrumentation, monitoring configuration, and data validation.
  • Automate alerting and remediation workflows to improve mean time to resolution and service uptime.
  • Establish observability standards with application, cloud, network, and security teams.
  • Create documentation, operational guidance, and best practices; mentor team members and support technical outcomes.
  • Participate in an on-call rotation for observability platform health and incident response.

Requirements

  • Bachelor’s degree in Computer Science or a related field, or relevant experience in lieu of a degree.
  • 5–7 years of experience as a systems engineer or in a related technical engineering role.
  • Hands-on Dynatrace experience covering application performance monitoring, infrastructure monitoring, real user monitoring, and Dynatrace Query Language.
  • Experience with Cisco ThousandEyes for synthetic testing, path visualization, and internet or wide area network performance monitoring.
  • Experience using Terraform and the Dynatrace Terraform provider for configuration-as-code.
  • Experience with Azure DevOps source control and pipelines; familiarity with GitHub Enterprise is helpful.
  • Strong Python and JavaScript programming skills for automation, custom telemetry, and tooling.
  • Experience developing custom Dynatrace extensions and implementing observability for Azure DevOps and CI/CD pipelines.
  • Foundational knowledge of applications, servers, storage, and networks, plus strong troubleshooting skills across metrics, logs, and traces.
  • Clear communication, cross-team collaboration, vendor coordination, and mentoring abilities.
  • Preferred: PowerShell, Splunk, Microsoft SCOM, SiteScope, Nagios, Monaco, Dynatrace Grail, Dynatrace or Azure certifications, and familiarity with Site Reliability Engineering practices.

Benefits

  • Full-time Monday–Friday schedule during Eastern business hours with 40 scheduled weekly hours.
  • Hybrid work model with on-site work at local Mass General Brigham sites weekly or monthly as business needs require.
  • Required on-call rotation, typically one week at a time, supporting platform health and incident response.
  • Periodic in-person stakeholder and team meetings; remote work requires a stable, secure, compliant workstation and Microsoft Teams participation using MGB-provided equipment.
  • Comprehensive benefits, career advancement opportunities, differentials, premiums, bonuses as applicable, and recognition programs.

Tech Stack

AzureJavaScriptNagiosPowerShellPythonSplunkTerraform

Categories

DevOpsSite Reliability
Mass General Brigham

About Mass General Brigham

10,000+ employees
Contact me