Microsoft

Principal Software Engineer, ML & Distributed Systems

Microsoft
Apply
1 day ago
Redmond, WA, USAStaff+
H1B Sponsor

Base Salary

$143k - $275k/yr

Responsibilities

  • Design, build, and operationalize scalable machine learning and deep learning models using containers and orchestration platforms.
  • Develop LLM prompt and fine-tuning strategies, evaluation pipelines, and optimizations for model quality, latency, and cost.
  • Architect, implement, and operate highly available, fault-tolerant, low-latency distributed services at hyperscale.
  • Build model serving and inference infrastructure with caching, batching, GPU capacity management, and large-scale A/B experimentation.
  • Design scalable APIs, data pipelines, and feature or signal stores for secure and reliable data flow between ML systems and product surfaces.
  • Drive instrumentation, monitoring, capacity planning, incident response, and live-site excellence for ML-backed services.
  • Collaborate with applied scientists, data scientists, backend engineers, and product teams to deliver production ML systems.
  • Participate in code reviews and architectural discussions and mentor engineers across ML and systems disciplines.

Requirements

  • Bachelor’s degree in computer science or a related technical field and 6+ years of technical engineering experience with coding in C, C++, C#, Java, JavaScript, or Python, or equivalent experience.
  • Preferred qualifications include a master’s degree and 8+ years of technical engineering experience, or a bachelor’s degree and 12+ years of technical engineering experience, or equivalent experience.
  • Proven experience designing, developing, and operating multi-tiered distributed services at scale.
  • Hands-on experience building, deploying, and operating production machine learning systems, including model training, evaluation, or LLM-based applications.
  • Experience across full product cycles from initial design through final delivery.
  • Experience with prompt engineering, retrieval-augmented generation, fine-tuning, and model evaluation frameworks.
  • Experience with distributed training, inference optimization, GPU capacity management, Kubernetes, large-scale data systems, streaming, Redis, feature stores, and experimentation platforms.
  • Ability to work across the ML and distributed systems boundary.

Benefits

  • Employees within the stated commute range of designated offices are expected to work from the office at least four days per week beginning January 26, 2026, subject to local law and jurisdiction.
  • Certain roles may be eligible for benefits and other compensation.
Microsoft

About Microsoft

10,000+ employees

Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.

Contact me