11 hours ago
Remote, IndiaSenior
Responsibilities
- Engineer and maintain the Dynatrace multi-tenant environment, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries.
- Own alert lifecycle standards covering ownership, naming, severity mapping, disposition taxonomies, production readiness, tuning, suppression, correlation, and retirement.
- Reduce operational alert fatigue while maintaining reliable detection of real client-impacting issues.
- Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths.
- Create standardized DQL queries, notebooks, dashboards, reporting baselines, and reliability metrics for operations, leadership, and service-level accountability.
- Manage observability configurations using versioned, peer-reviewed, reproducible configuration-as-code practices.
- Enhance alert-to-ticket enrichment and automated remediation integrations with enterprise ITSM systems.
- Partner with platform, OS, SQL, API, automation, tools, and SRE specialists to resolve monitoring blind spots and validate runbooks.
- Maintain observability standards, standard operating procedures, reference materials, and Confluence documentation.
- Collaborate across operations, SRE, database, engineering, and leadership teams to translate infrastructure signals into service and business impacts.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related technical field, or an equivalent combination of education and experience.
- At least 7 years of experience in observability, application performance monitoring, monitoring engineering, site reliability engineering, or related technology infrastructure.
- Hands-on Dynatrace expertise, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric and custom events, and synthetic monitors.
- Demonstrated experience designing, configuring, and tuning enterprise-scale alerting systems to reduce noise while preserving detection capabilities.
- Experience integrating monitoring systems with enterprise ITSM or ticketing systems, such as Salesforce, for automated routing and enrichment.
- Experience in regulated hosting environments such as healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001.
- Dynatrace certification, or commitment to obtain certification within one year of hire.
- Strong knowledge of AWS infrastructure, Windows and Linux operating systems, SQL Server database concepts, cloud-native observability, systems integration, database operations, and scripting with Python or PowerShell.
- Strong problem-solving, written and verbal communication, organization, attention to detail, ownership, and cross-functional collaboration skills.
About NextGen
NEXTGEN is an Australia-based technology services and value-added distribution company that helps vendors and channel partners sell cybersecurity, cloud, enterprise software, and data management solutions. It offers software licensing, compliance and audit services, data centre and storage solutions, and go-to-market and digital marketing support. Founded in 2011 and headquartered in North Sydney, it is now part of Exclusive Networks, expanding its reach across international markets.
