Base Salary
$120k - $261k/yr
Responsibilities
- Own day-to-day operations for supercomputing clusters, GPU compute, and interconnect fabrics while ensuring GPU availability, service reliability, and AI training stability.
- Lead incident triage, mitigation, recovery, and root cause analysis for production issues across large-scale AI infrastructure.
- Debug failures across hardware provisioning, GPU interconnect fabrics, PCIe subsystems, and GPU interactions.
- Identify systemic failure patterns and create technical guidance, troubleshooting procedures, playbooks, and escalation frameworks.
- Design and use automation, telemetry, and diagnostic tooling to improve issue detection, observability, debuggability, and mean time to mitigation.
- Partner with internal engineering teams and manufacturers to drive operational and product fixes.
Requirements
- Bachelor’s degree in Computer Science or a related technical field and 4+ years of technical engineering experience with coding, or equivalent experience.
- Coding experience in one or more of C, C++, C#, Java, JavaScript, or Python.
- Ability to meet Microsoft Cloud Background Check and applicable customer or government security screening requirements.
- Preferred: master’s degree with 6+ years of experience, or bachelor’s degree with 8+ years of experience, or equivalent experience.
- Preferred: 4+ years operating HPC, AI, or large-scale distributed systems in production.
- Preferred: 1+ year operating interconnect fabrics for HPC, AI, or large-scale distributed systems in production.
- Preferred: Linux systems knowledge and experience debugging low-level infrastructure issues.
- Preferred: ability to diagnose production issues across hardware, firmware, drivers, and software stacks.
About Microsoft
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.
