18 hours ago
Base Salary
$143k - $304k/yr
Responsibilities
- Design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring.
- Define technical strategy and architecture for New Product Introduction, including platform bring-up, validation, deployment readiness, and lifecycle management.
- Lead hardware and firmware reliability architecture involving health telemetry, diagnostics, failure detection, predictive insights, and remediation automation.
- Drive failure analysis and root-cause investigations for critical live-site incidents across hardware, firmware, OS, storage, networking, and infrastructure platforms.
- Partner with Microsoft 365, Azure Hardware Systems, silicon vendors, OEM/ODM partners, firmware teams, and operations organizations.
- Establish engineering standards, observability patterns, and operational mechanisms that improve fleet availability and reduce deployment risk.
- Evolve Microsoft 365 substrate infrastructure strategy for AI and agentic workloads across compute, memory, storage, and networking.
- Influence cross-organizational technical direction, communicate architectural tradeoffs, and mentor early-career engineers.
Requirements
- Bachelor’s degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding, or equivalent experience.
- Experience with coding in C, C++, C#, Java, JavaScript, Python, or comparable languages.
- Ability to meet Microsoft, customer, and/or government security screening requirements, including the Microsoft Cloud background check.
- Preferred: master’s degree with 8+ years of experience, or bachelor’s degree with 12+ years of experience, or equivalent experience.
- Experience leading architecture and technical strategy for large-scale distributed systems, cloud infrastructure, or hyperscale service platforms.
- Deep experience with hardware platforms, firmware, operating systems, drivers, datacenter infrastructure software, storage systems, networking, or platform software.
- Experience with New Product Introduction, platform validation, deployment readiness, lifecycle management, or fleet-scale hardware operations.
- Experience using telemetry, observability, failure analysis, predictive diagnostics, or automation to improve hyperscale infrastructure reliability.
- Experience working with silicon providers, hardware vendors, OEMs, ODMs, firmware teams, or platform engineering organizations.
- Experience influencing cross-organizational engineering strategy and driving complex technical programs across multiple teams.
- Ability to explain technical tradeoffs to senior engineering and business leaders; experience with Microsoft Secure environments or US Government clouds is preferred.
Benefits
- The role is based in Redmond, Washington and requires working in the office at least three days per week.
- Certain roles may be eligible for benefits and other compensation; additional benefits and pay information is provided through Microsoft’s careers site.
About Microsoft
Microsoft develops operating systems, productivity software, cloud services, developer tools, and consumer devices for individuals, enterprises, and governments. Its main products include Windows, Microsoft 365, Azure, Visual Studio/GitHub, Xbox, and LinkedIn; revenue comes from software subscriptions and licenses, cloud consumption, hardware sales, and advertising. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is a public company traded on Nasdaq.
