Base Salary
$102k - $202k/yr
Responsibilities
- Lead high-severity incidents across Azure services as the accountable incident commander, directing detection, triage, resolution, and customer communications.
- Coordinate real-time decisions across Engineering, Support, Product Management, Communications, Field teams, and external vendors during live-site incidents.
- Perform production triage, root-cause analysis, product-gap identification, incident reviews, and preventative service improvements.
- Design and promote resilient cloud infrastructure architectures, operational frameworks, mitigation playbooks, auto-remediation tools, and customer self-service capabilities.
- Use telemetry, support cases, feedback, and platform health metrics to improve availability, reliability, observability, and supportability.
- Participate in the 24/7 on-call rotation and build cross-functional partnerships across engineering, business, and support organizations.
Requirements
- Bachelor’s degree in a relevant field and 2+ years of technical experience in software engineering, network engineering, service engineering, systems engineering, or industrial controls, or equivalent experience.
- 2–4+ years of experience in cloud operations, incident response, SRE, or large-scale systems engineering, preferably with Azure, AWS, or GCP.
- Service Engineering experience in a 24/7/365 enterprise environment.
- Strong incident command, crisis management, communication, judgment, negotiation, decision-making, and problem-resolution skills under pressure.
- Deep understanding of cloud architecture patterns, microservices, containerization, high availability, disaster recovery, business continuity, and performance tuning.
- Familiarity with monitoring and observability tools and fluency in one or more automation languages such as PowerShell, Python, or CLI.
- Understanding of ITIL or other incident management frameworks and the ability to diagnose and debug user code on Windows or Linux platforms.
- Preferred qualifications include 4+ years as an Incident Commander or Crisis Manager, SRE practices, chaos engineering, fault injection, AI/ML production concepts, and relevant cloud or ITIL/SRE certifications.
- Ability to meet Microsoft Cloud and applicable customer or government security screening requirements.
Benefits
- The role includes participation in an on-call rotation for live-site incident response.
- The position is open to U.S. applicants, with location-specific base-pay ranges for the San Francisco Bay Area and New York City metropolitan area.
- Certain roles may be eligible for benefits and additional compensation.
- Applications are accepted on an ongoing basis until the position is filled, with the posting open for a minimum of five days.
Tech Stack
Categories
About Microsoft
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.
