22 hours ago
Base Salary
$151k - $235k/yr
Responsibilities
- Build and own large-scale automation infrastructure for accelerator fleet health, including fault detection, diagnosis, triage, and remediation.
- Design test frameworks, diagnostic tooling, coverage strategies, and qualification automation for hardware bring-up, regression, and production.
- Develop predictive failure detection using telemetry, sensor data, error trends, and log correlation.
- Create monitoring dashboards, alerting, and fleet-health metrics for manufacturing, lab, and production environments.
- Debug system-level issues across GPU, compute, networking, Linux, PCIe, power, NIC, NVMe, and accelerator subsystems on x86 and ARM.
- Perform root-cause analysis across firmware, kernels, drivers, and physical-layer systems and drive permanent corrective actions.
- Develop Linux device drivers, OS-level software, data pipelines, and CI/CD pipelines for manufacturing and production deployments.
- Collaborate with engineering teams, ODMs, design partners, and datacenter operations on hardware readiness, testability, diagnostics, and field-failure improvements.
- Participate in post-incident reviews and maintain documentation, runbooks, and automation.
Requirements
- At least 6 years of experience in systems design, software development, operations, automation, and process improvement.
- At least 6 years of non-internship professional software development experience.
- At least 6 years of programming experience with one or more modern languages such as C++, C#, Java, Python, Golang, PowerShell, or Ruby.
- At least 6 years of designing or architecting reliable and scalable systems.
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experience.
- At least 5 years of Linux/Unix operating systems experience.
- Preferred qualifications include a master’s degree or higher, 5+ years of full software development lifecycle experience, large-scale computing automation experience, hardware troubleshooting or CUDA/kernel experience, hardware design and validation experience, and experience leading technical teams or working in data center engineering or operations.
Benefits
- Comprehensive health insurance including medical, dental, vision, prescription, life, AD&D, EAP, mental health support, medical advice, flexible spending accounts, and adoption and surrogacy reimbursement coverage.
- 401(k) matching, paid time off, parental leave, sign-on payments, and restricted stock units.
- The role is based in Cupertino, Austin, or Seattle and may require occasional regional and international travel of less than 10% to design and manufacturing partner sites.
About Amazon
Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.
