Systems Engineer, Senior - Observability
$93,954–$136,739 year
On-siteSomerville, Massachusetts, United States
Job Summary
Deploy, configure, and optimize enterprise observability platforms, primarily Dynatrace, alongside Cisco ThousandEyes for network-path and digital experience monitoring. Build dashboards, alerts, and synthetic tests to enable faster issue detection, while developing Terraform configuration-as-code solutions and custom Dynatrace extensions. Implement observability for DevOps frameworks in Azure, including instrumentation and quality gates within CI/CD pipelines, and automate alerting workflows to reduce mean time to resolution. Partner with application, cloud, network, and security teams to establish monitoring standards, create operational documentation, and onboard infrastructure to observability platforms. Work full-time in a hybrid model with an on-call rotation, supporting the Digital Enterprise Observability team's mission to strengthen reliability across Mass General Brigham's critical digital services.
Required Qualifications
- Bachelor's degree in Computer Science or a related field
- 5–7 years of experience as a systems engineer or in a related technical engineering role
- Hands-on experience with Dynatrace, including application performance monitoring, infrastructure monitoring, real user monitoring, and Dynatrace Query Language
- Experience with Cisco ThousandEyes for synthetic testing, path visualization, and internet or wide area network performance monitoring
- Experience using Terraform for configuration as code, including the Dynatrace Terraform provider
- Experience with source control and pipelines in Azure DevOps
- Strong programming skills in Python and JavaScript for automation, custom telemetry, and tooling
- Experience developing custom Dynatrace extensions
- Experience implementing observability for DevOps frameworks and CI/CD pipelines in Azure
- Foundational knowledge of technology infrastructure, including applications, servers, storage, and networks
- Strong analytical and troubleshooting skills across metrics, logs, and traces
- Ability to communicate clearly with both technical and non-technical stakeholders
- Strong collaboration skills, including experience working across teams and with vendors to achieve results
- Ability to mentor team members and support better technical outcomes
- Full-time schedule, Monday through Friday during Eastern business hours
- Participation in an on-call rotation, typically one week at a time, to support observability platform health and incident response
- Hybrid work model with on-site work at Mass General Brigham local sites on a weekly or monthly basis, depending on business needs
- Availability for periodic in-person stakeholder meetings, team meetings, and internal customer needs
- On remote workdays, employees must work from a stable, secure, and compliant workstation in a quiet environment
- Microsoft Teams video participation is required using MGB-provided equipment
Desired Qualifications
- Relevant experience may be considered in lieu of a degree
- Familiarity with GitHub Enterprise
- PowerShell scripting experience
- Familiarity with Splunk, Microsoft SCOM, SiteScope, or Nagios
- Dynatrace Associate or Professional certification
- Azure certification, such as Azure Administrator or Azure DevOps Engineer
- Experience with Monaco and Dynatrace Grail
- Exposure to Site Reliability Engineering practices, including service level indicators, service level objectives, and error budgets
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.