16 hours ago
Base Salary
$120k - $261k/yr
Responsibilities
- Design and build large-scale distributed services for fleet reliability, hardware health, and operational efficiency.
- Develop telemetry and analytics platforms that process infrastructure health signals at hyperscale.
- Build predictive models and intelligent services for hardware failure detection, repair recommendations, anomaly detection, and fleet risk forecasting.
- Analyze telemetry from servers, storage, networking, rack infrastructure, and datacenter systems to improve reliability and availability.
- Develop AI-assisted experiences for incident investigation, root cause analysis, and repair decision-making.
- Design safe automation and remediation workflows while maintaining strong operational controls.
- Partner with hardware, reliability, capacity planning, Fleet Management, and Azure Infrastructure teams.
- Participate in architecture reviews, code reviews, and live-site operations.
- Drive projects from design through deployment and operational ownership, mentor engineers, and contribute to engineering excellence.
Requirements
- Bachelor’s degree in Computer Science or a related technical field and 4+ years of technical engineering experience with coding, or equivalent experience.
- Proficiency with one or more of C, C++, C#, Java, JavaScript, or Python.
- Ability to meet Microsoft, customer, and/or government security screening requirements, including the Microsoft Cloud background check.
- Experience designing and operating distributed systems and cloud services at scale is preferred.
- Experience with hardware infrastructure, storage systems, server platforms, networking systems, or datacenter operations is preferred.
- Experience with data science, statistics, machine learning, forecasting, anomaly detection, or predictive analytics is preferred.
- Experience with telemetry and data platforms such as Azure Data Explorer (Kusto), Spark, Fabric, or Databricks is preferred.
- Experience developing AI-powered operational tools, intelligent automation systems, or agent-based solutions is preferred.
- Experience with hardware reliability engineering, fleet management, capacity planning, or infrastructure health monitoring is preferred.
- Experience with Microsoft 365 components such as Exchange, Substrate, or SharePoint is preferred.
- Ability to independently drive complex technical projects from concept through production deployment and influence across organizations is preferred.
- Tier 2 or Tier 3 United States Government clearance to work in secure Microsoft cloud environments is preferred.
- A master’s degree with 6+ years of experience or a bachelor’s degree with 8+ years of experience is also listed as a preferred qualification, or equivalent experience.
Benefits
- The position is based in Redmond, Washington and requires working in the office at least 3 days per week.
- Certain roles may be eligible for benefits and other compensation.
Tech Stack
Categories
About Microsoft
Microsoft develops operating systems, productivity software, cloud services, developer tools, and consumer devices for individuals, enterprises, and governments. Its main products include Windows, Microsoft 365, Azure, Visual Studio/GitHub, Xbox, and LinkedIn; revenue comes from software subscriptions and licenses, cloud consumption, hardware sales, and advertising. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is a public company traded on Nasdaq.
