Base Salary
$143k - $275k/yr
Responsibilities
- Own InfiniBand and GPU interconnect fabric operations across large-scale AI supercomputing environments, ensuring GPU availability, training stability, and SLA compliance.
- Lead high-severity fabric incidents through detection, triage, mitigation, recovery, and root-cause analysis.
- Debug issues across InfiniBand, Subnet Manager, GPU interconnects, PCIe, GPUs, firmware, drivers, operating systems, and services.
- Define reliability models, failure domains, operational standards, playbooks, technical support guidance, and escalation frameworks.
- Architect automation, telemetry, diagnostics, and tooling to improve detection, observability, debuggability, and time to mitigation.
- Provide technical leadership, influence engineering direction, mentor senior engineers, and partner with platform, hardware, firmware, and service teams.
Requirements
- Bachelor’s degree in computer science or a related technical field and six or more years of technical engineering experience with coding, or equivalent experience.
- At least six years of experience operating large-scale distributed systems, high-performance computing, or artificial intelligence infrastructure in production environments.
- Demonstrated ownership of mission-critical production infrastructure affecting service availability, GPU workloads, and customer SLAs.
- Hands-on experience operating and debugging interconnect fabrics for large-scale compute workloads.
- Strong Linux systems knowledge and experience debugging low-level infrastructure issues across operating systems, drivers, and services.
- Ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve complex production issues.
- Ability to meet Microsoft Cloud Background Check and applicable customer or government security screening requirements.
- Preferred qualifications include a bachelor’s degree with ten or more years of experience, or a master’s degree with eight or more years of experience, or equivalent experience.
Benefits
- The role is a full-time Microsoft position with a U.S. base pay range of $142,800–$274,800 per year; location-specific ranges apply in the San Francisco Bay Area and New York City metropolitan area.
- Certain roles may be eligible for benefits and other compensation.
- Applications are accepted on an ongoing basis, with the position open for a minimum of five days.
About Microsoft
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.
