5 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect and develop distributed control-plane services for provisioning, orchestration, and lifecycle management of AI accelerator infrastructure.
- Design scalable systems for state management, health monitoring, reconciliation, and fault recovery.
- Build host and device management software connecting cloud infrastructure with operating systems, drivers, firmware, and accelerator devices.
- Develop infrastructure for accelerator virtualization, device assignment, isolation, and resource management.
- Integrate accelerator infrastructure with Kubernetes and cloud-native platforms.
- Design APIs and abstractions across cloud services, host software, and device interfaces.
- Improve reliability, security, observability, diagnostics, testing, and operational readiness.
- Diagnose complex issues spanning distributed services, operating systems, host software, drivers, firmware, and hardware.
- Lead architecture and design reviews, solve cross-layer system issues, mentor engineers, and raise technical standards.
- Champion AI-assisted engineering practices and workflows for architecture, coding, debugging, testing, documentation, and automation.
- Collaborate with hardware, firmware, operating system, cloud infrastructure, and AI platform teams on end-to-end solutions.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical discipline, or equivalent experience with at least 12+ years of relevant industry experience.
- Software engineering experience with C++, C, C#, Rust, Go, or similar languages.
- Experience designing and developing distributed systems, cloud infrastructure, or systems software.
- Understanding of concurrency, state management, asynchronous programming, and failure recovery.
- Experience building reliable production software with strong debugging and problem-solving skills.
- Demonstrated technical leadership across complex, multi-team engineering projects.
- Ability to use AI-assisted engineering tools and workflows with sound engineering judgment.
- Ability to develop expertise in complex technical domains and raise team competency through mentoring and knowledge sharing.
- Preferred experience with cloud control planes, resource orchestration, infrastructure management, Kubernetes, containers, and cloud-native infrastructure.
- Preferred experience with GPU, AI accelerator, heterogeneous compute, PCIe, SR-IOV, PF/VF, device virtualization, device passthrough, host/device agents, Linux systems software, drivers, or firmware interfaces.
- Preferred experience with hardware lifecycle management, health monitoring, fault recovery, telemetry, observability, and debugging large-scale production systems.
- Preferred experience applying AI-assisted software engineering and establishing responsible AI-enabled workflows for development, debugging, testing, documentation, and engineering automation.
About Microsoft
Microsoft develops operating systems, productivity software, cloud services, developer tools, and consumer devices for individuals, enterprises, and governments. Its main products include Windows, Microsoft 365, Azure, Visual Studio/GitHub, Xbox, and LinkedIn; revenue comes from software subscriptions and licenses, cloud consumption, hardware sales, and advertising. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is a public company traded on Nasdaq.
