7 days ago
Base Salary
$120k - $261k/yr
Responsibilities
- Define architecture for agentic reliability systems covering monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge.
- Build automated capabilities for detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning.
- Establish standards for identity, access, compliance, rollback, auditability, change management, and human escalation in agentic operations.
- Drive adoption of monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices across service teams.
- Lead live-site investigations and convert reliability gaps and incident patterns into platform investments and reusable automation.
- Mentor engineers and shape technical direction across monitoring, observability, incident response, and AI-assisted automation.
Requirements
- Bachelor’s degree in computer science or a related technical field and 4+ years of technical engineering experience with coding in C, C++, C#, Java, JavaScript, or Python, or equivalent experience.
- Ability to pass Microsoft Cloud Background Check screening upon hire or transfer and every two years thereafter.
- Preferred experience building production-scale platforms, cloud services, distributed systems, or reliability automation.
- Preferred experience architecting systems across service boundaries and driving execution across partner teams without direct authority.
- Preferred experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging.
- Preferred judgment regarding production safety, automation risk, customer impact, security, privacy, compliance, and responsible AI-assisted and agentic automation.
- Preferred experience with agentic automation, AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platforms.
- Preferred experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, GitHub, or similar ecosystems.
- Preferred experience creating reliability metrics such as SLO compliance, alert quality, time to detect, time to mitigate, automation coverage, and incident recurrence.
- Preferred experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering.
Benefits
- The typical U.S. base pay range is USD $119,800–$234,700 per year; a separate range applies in the San Francisco Bay Area and New York City metropolitan area.
- Certain roles may be eligible for benefits and other compensation.
- The position will remain open for at least five days and applications are accepted on an ongoing basis until filled.
Tech Stack
Categories
Site Reliability
About Microsoft
Microsoft develops operating systems, productivity software, cloud services, developer tools, and consumer devices for individuals, enterprises, and governments. Its main products include Windows, Microsoft 365, Azure, Visual Studio/GitHub, Xbox, and LinkedIn; revenue comes from software subscriptions and licenses, cloud consumption, hardware sales, and advertising. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is a public company traded on Nasdaq.
