
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
RaceTrac, Inc.8 days ago
Atlanta, GA, USASenior
Responsibilities
- Serve as the primary Dynatrace and observability subject-matter expert across the organization.
- Design, develop, and optimize enterprise observability solutions, including DQL queries, dashboards, workflows, alerts, analytics, and telemetry standards.
- Build and maintain monitoring for Azure services, APIs, integrations, backend systems, distributed applications, and mobile platforms.
- Analyze logs, metrics, traces, telemetry, and distributed transactions to identify root causes, latency issues, performance bottlenecks, and reliability risks.
- Enable observability, tracing, crash analytics, and end-user experience monitoring for iOS and Android applications.
- Read and analyze .NET application code to troubleshoot application behavior, performance, deployment, and runtime issues.
- Support Azure Function deployments, configuration, scaling, and troubleshooting, and ensure observability readiness before production releases.
- Conduct root-cause analysis, incident response, blameless postmortems, and preventive reliability improvements.
- Automate operational workflows and monitoring processes while reducing alert noise, manual intervention, recurring incidents, and resolution time.
- Support large-scale services throughout their lifecycle, including design, deployment, operation, capacity planning, launch reviews, and production support.
- Collaborate with engineering, cloud, mobile, SRE, operations, and business teams on architecture reviews, releases, and operational excellence.
- Improve and tune operational efficiency in Windows-based production infrastructure and support applications across private and public cloud environments.
Requirements
- 10+ years of overall IT experience.
- At least 4 years of working experience with Azure.
- Expert-level hands-on experience with Dynatrace and advanced experience with Dynatrace Query Language (DQL).
- Strong hands-on expertise with Azure Kusto Query Language (KQL), telemetry, observability, distributed tracing, metrics, and logging.
- Deep experience with Azure Monitor, Application Insights, Azure Functions, Azure API Management, Azure Log Analytics, and App Services.
- Strong understanding of API architectures, API gateways, backend integrations, distributed systems, and enterprise application architectures.
- Prior hands-on experience developing .NET applications and the ability to read, analyze, and understand .NET code.
- Experience troubleshooting and deploying Azure Functions and cloud-native applications.
- Experience enabling observability and telemetry for iOS and Android applications, including mobile telemetry, crash analytics, API monitoring, and end-user experience monitoring.
- Experience programming in at least one of C, C++, Java, Python, or Go.
- Experience with Jenkins or a similar application.
- General knowledge of infrastructure-as-code and configuration-management tools such as Terraform, Ansible, Chef, Puppet, or SCCM.
- Comfort working with large-scale production systems, load balancing, monitoring, distributed systems, and configuration management.
- Ability to design, analyze, troubleshoot, debug, optimize code, and automate routine tasks.
- Preferred experience with OpenTelemetry implementation and instrumentation, AI-driven observability, AIOps, high-volume enterprise digital platforms, ServiceNow, Databricks, SQL platforms, and integration technologies.
- A bachelor’s degree in Computer Science or a related field is preferred; equivalent practical experience will be considered.
- Strong analytical, troubleshooting, communication, stakeholder-management, collaboration, and proactive problem-solving skills.
Benefits
- RaceTrac describes opportunities for growth and new career paths across its operating divisions.
- The role collaborates across corporate, engineering, cloud, mobile, operations, and business teams in private- and public-cloud environments.
- The employer states that all qualified applicants will receive consideration without regard to protected characteristics.
Categories
Site Reliability