Amazon

SDE II, ML Infra Services, Annapurna Labs

Amazon
Apply
1 day ago
Seattle, WA, USAMid Level
H1B sponsor

Base Salary

$144k - $194k/yr

Responsibilities

  • Lead the design and implementation of ML infrastructure for capacity management, workload scheduling, and fleet orchestration across ML accelerators.
  • Build systems that provide low-wait-time, high-utilization, zero-configuration access to ML accelerators from various environments.
  • Design and code scalable solutions, create metrics, implement automation, and resolve the root causes of software defects.
  • Participate in architecture and design discussions, code reviews, and technical communication with internal and external stakeholders.
  • Collaborate cross-functionally with ML scientists, training infrastructure engineers, hardware teams, and internal customers.
  • Deliver high-impact solutions for a large customer base and drive architectural decisions in a startup-like environment.

Requirements

  • At least 3 years of non-internship professional software development experience.
  • At least 2 years of non-internship experience designing or architecting new and existing systems, including design patterns, reliability, and scaling.
  • Experience programming with at least one software programming language.
  • Preferred: 3+ years of full software development lifecycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Preferred: 2+ years building large-scale machine-learning infrastructure for recommendation, ads ranking, personalization, or search.
  • Preferred: strong proficiency in Go or Java, Python, and working knowledge of JavaScript or TypeScript.
  • Preferred: experience building and operating large-scale distributed systems on Kubernetes.
  • Preferred: experience designing, deploying, and maintaining production services at scale, including on-call ownership.
  • Preferred: experience with ML infrastructure orchestration, scheduling, or resource management at scale.
  • Preferred: proficiency in application- and kernel-level performance profiling and optimization, integrated software/hardware performance analysis, and observability and telemetry.
  • Preferred: experience debugging complex large-scale distributed systems and driving best practices.

Benefits

  • Comprehensive benefits including medical, dental, vision, prescription, life and AD&D insurance, employee assistance, mental health support, medical advice, flexible spending accounts, adoption and surrogacy reimbursement, 401(k) matching, paid time off, and parental leave.
  • The team supports workplace flexibility and work/life balance.
  • Employees receive mentorship, knowledge-sharing, and one-on-one code reviews.
  • The role is based in Seattle, Washington, USA.
Amazon

About Amazon

10,000+ employees

Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.

Contact me