Amazon

System Development Engineer II, HWEngS Ultraservers

Amazon
Apply
1 day ago
Cupertino, CA, USA +2 moreMid Level
H1B sponsor

Base Salary

$129k - $201k/yr

Responsibilities

  • Build and own large-scale automation infrastructure for accelerator fleet health, including fault detection, diagnosis, triage, and remediation.
  • Design test frameworks, diagnostic tools, test coverage strategies, and qualification automation for hardware bring-up, regression, and production.
  • Develop predictive failure detection using telemetry, sensor data, error trends, and log correlation.
  • Create monitoring dashboards, alerting, and fleet-health metrics for manufacturing, lab, and production environments.
  • Debug system-level failures across GPUs, networking, Linux boot and runtime, PCIe, power, NIC, NVMe, firmware, kernels, drivers, and physical-layer components on x86 and ARM.
  • Build data pipelines correlating test results, sensor telemetry, and component data to identify systemic yield issues.
  • Develop and maintain Linux device drivers and work with OS internals and accelerator/GPU software stacks.
  • Build and manage system tests and deployment pipelines for manufacturing lines and production fleets.
  • Collaborate with engineering teams, internal customers, ODMs, design partners, and datacenter operations on hardware readiness, testability, and design improvements.
  • Participate in incident reviews, drive permanent corrective actions, and maintain documentation, runbooks, and automation.

Requirements

  • At least 2 years of non-internship professional software development experience.
  • At least 2 years of experience designing or architecting reliable and scalable systems.
  • At least 2 years of networking, storage systems, operating systems administration, and hands-on systems engineering experience.
  • Programming experience in at least one of C++, C#, Java, Python, Golang, PowerShell, or Ruby.
  • At least 2 years of Linux operating-systems experience.
  • Experience leading the design, development, and deployment of complex, performant, reliable, and scalable production software solutions.
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experience.
  • Preferred: master's degree in computer science, electrical engineering, or a related field.
  • Preferred: at least 2 years of full software development lifecycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Preferred: at least 2 years of Linux kernel driver development, OS internals, or GPU/accelerator driver integration and diagnostics.
  • Preferred: experience troubleshooting low-level system integration issues involving server and GPU hardware, BMC/IPMI, firmware, PCIe topology, and hardware-level fault isolation.

Benefits

  • Comprehensive health coverage including medical, dental, vision, prescription, life and AD&D insurance, an employee assistance program, mental health support, a medical advice line, and flexible spending accounts.
  • Adoption and surrogacy reimbursement coverage, 401(k) matching, paid time off, and parental leave.
  • Sign-on payments and restricted stock units may be included in the overall Amazon package; final compensation depends on experience, qualifications, and location.
  • Role is based in Cupertino, Austin, or Seattle, with occasional regional and international travel of less than 10% to design and manufacturing partner sites.

Tech Stack

AWSC#C++GoJavaLinuxPowerShellPythonRuby

Categories

Amazon

About Amazon

10,000+ employees

Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.

Contact me