14 hours ago
Base Salary
$120k - $261k/yr
Responsibilities
- Create software services, applications, and tools that are secure, highly available, scalable, and reliable.
- Analyze live and historical telemetry to build and maintain fleet-wide datacenter risk tooling.
- Investigate and apply predictive risk models, forecasting, and agentic triage to improve operational insights and reduce manual intervention.
- Lead architecture discussions, create solution proposals, and refine implementation plans through design evaluation.
- Automate assessment and reporting of manually identified critical environment risks.
- Optimize, debug, refactor, and reuse code to improve performance, maintainability, effectiveness, and return on investment.
- Apply metrics, coding patterns, and best practices to improve code quality and stability.
- Collaborate with software development and data engineering teams on security, privacy, and reliability.
- Use AI-assisted engineering tools across code authoring, review, and test generation and help establish responsible team practices.
- Mentor mid-level engineers and onboard new engineers and interns through design reviews, pair programming, code reviews, and operational guidance.
Requirements
- Bachelor's degree in computer science or a related technical field and at least four years of technical engineering experience with coding, or equivalent experience.
- Six or more years of end-to-end software engineering experience, or six or more years as a software engineer, database engineer, site reliability engineer, or similar role.
- Experience with programming languages including C, C++, C#, Java, JavaScript, or Python.
- Experience with modern DevOps practices, build and deployment pipelines, YAML development, and unit or integration testing.
- Production experience with Python/PySpark, TypeScript, Kusto, or another query language.
- Hands-on experience with Azure Synapse Analytics, Azure Log Analytics, and analytical or visualization tools such as Web Apps or Power BI.
- Hands-on experience applying AI-assisted development tools and AI/ML or LLM-based techniques to production systems.
- At least four years of experience building or analyzing risk, health, or reliability tooling for datacenter critical environments, information technology, or related environments and developing automated monitoring.
- Demonstrated experience mentoring engineers or leading interns and growing technical capability within a team.
- Ability to collaborate with diverse global teams, work independently, communicate effectively with words and data, and travel as necessary to support operations in the United States.
- Ability to meet Microsoft Cloud background check requirements upon hire or transfer and every two years thereafter.
Benefits
- The position is open for a minimum of five days, with applications accepted on an ongoing basis until filled.
- Certain roles may be eligible for benefits and other compensation.
- Equal opportunity employment and religious or disability accommodation support are provided.
Tech Stack
Categories
About Microsoft
Microsoft develops operating systems, productivity software, cloud services, developer tools, and consumer devices for individuals, enterprises, and governments. Its main products include Windows, Microsoft 365, Azure, Visual Studio/GitHub, Xbox, and LinkedIn; revenue comes from software subscriptions and licenses, cloud consumption, hardware sales, and advertising. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is a public company traded on Nasdaq.
