
Sr. Observability Engineer
FreedomPay1 month ago
Remote, United StatesSenior
Responsibilities
- Define and evolve the enterprise observability vision, standards, principles, and roadmap.
- Lead implementation, optimization, and adoption of Dynatrace SaaS and its platform capabilities.
- Establish scalable telemetry architecture and OpenTelemetry standards.
- Design observability solutions for Kubernetes, containers, microservices, distributed applications, and public cloud environments.
- Apply AIOps and agentic operations for detection, event correlation, diagnosis, automation, operational response, and continuous improvement.
- Develop service health models, SLOs, dashboards, alerting strategies, telemetry governance, and observability best practices.
- Collaborate with engineering, site reliability, platform, security, infrastructure, operations, and architecture teams.
- Guide migration from legacy monitoring to modern observability models and enterprise platform transformation.
- Mentor engineers, influence technical direction, solve complex problems, and shape architecture and engineering standards.
Requirements
- Master’s degree from an accredited college or university in Computer Science, Information Systems, Engineering, or a related technical field.
- 8+ years of experience in observability, monitoring, site reliability engineering, platform engineering, infrastructure engineering, or related technical disciplines.
- Deep hands-on experience with Dynatrace SaaS and recent Dynatrace platform capabilities and innovations.
- Strong experience with OpenTelemetry, telemetry instrumentation, collection, and architecture practices.
- Experience designing and implementing observability solutions for cloud-native infrastructure, distributed systems, and modern application environments.
- Hands-on experience implementing AIOps and/or agentic operations capabilities.
- Strong understanding of metrics, logs, traces, event correlation, service health, alerting models, and operational intelligence.
- Ability to lead technical initiatives, influence architectural decisions, and drive adoption across teams and stakeholders.
- Strong verbal and written communication skills.
- Preferred: experience with enterprise-scale observability transformation initiatives; AWS, Azure, and/or Google Cloud Platform; SRE principles including SLIs, SLOs, incident response, and operational resilience; automation, orchestration, and remediation workflows; additional observability platforms or ecosystems; and Python, Go, Bash, or JavaScript.