Base Salary
$143k - $275k/yr
Responsibilities
- Collaborate with engineers and researchers to build and optimize training infrastructure and tools for LLMs, SLMs, multimodal models, and code-specific models.
- Design, build, and improve highly scalable and reliable production services, APIs, backend systems, and control-plane systems.
- Implement services that support production traffic while meeting security and privacy requirements.
- Improve engineering systems and practices for service quality in complex cloud environments.
- Deploy, monitor, and improve services in production.
- Lead technical design and retrospective efforts, mentor engineers, and drive ambiguous projects to completion.
Requirements
- Bachelor’s degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding, or equivalent experience.
- Preferred qualifications include a master’s degree and 8+ years of experience, or a bachelor’s degree and 12+ years of experience, or equivalent experience.
- At least 5 years of software engineering experience with significant ownership of production services, cloud platforms, distributed systems, or developer infrastructure.
- Strong experience building and operating containerized platforms with Kubernetes or similar orchestration systems.
- Strong coding skills in one or more systems or backend languages, including Python, Go, Rust, C++, C#, or Java.
- Experience designing reliable production APIs, backend services, or control-plane systems managing compute, storage, networking, or runtime environments.
- Understanding of cloud infrastructure fundamentals including identity, networking, storage, observability, capacity planning, security, and safe deployment practices.
- Experience diagnosing production issues using logs, metrics, traces, dashboards, and incident response processes.
- Experience with Microsoft Azure, AWS, or Google Cloud and managed cloud services such as Kubernetes, container registries, object storage, private networking, identity, secrets, and monitoring.
- Experience with multi-tenant platforms, sandboxed execution, remote development or hosted notebook environments, evaluation infrastructure, ephemeral compute, container image systems, caching, artifact distribution, and startup-latency optimization.
- Experience with cloud networking, secure runtime design, AI infrastructure, agent execution, evaluation platforms, GPU workloads, Windows/Linux environments, or VM/container hybrid systems.
- Ability to lead technical design and retrospectives, mentor engineers, collaborate across teams, and meet Microsoft Cloud Background Check and applicable customer or government security screening requirements.
Benefits
- The typical U.S. base pay range is $142,800–$274,800 annually, with a separate San Francisco Bay Area and New York City metropolitan area range.
- Certain roles may be eligible for benefits and other compensation.
- Applications are accepted on an ongoing basis until the position is filled, with the position open for a minimum of 5 days.
Tech Stack
Categories
About Microsoft
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.
