1 day ago
Base Salary
$106k - $155k/yr
Responsibilities
- Design, develop, and optimize HPC software for distributed and parallel workloads on large-scale Linux clusters.
- Optimize throughput, latency, scaling behavior, and power utilization across CPU, memory, storage, and network subsystems.
- Develop and maintain cluster bring-up, diagnostics, monitoring, power-usage, and health-check tooling.
- Translate workload characteristics into power-efficient HPC software solutions with algorithms, systems, and application teams.
- Collaborate on HPC node, storage, interconnect, CPU/GPU, memory, PCIe, NUMA, and network-topology requirements.
- Participate in hardware/software co-debugging, performance analysis, stability investigations, and failure analysis.
- Support rack-level integration involving power, cooling, cabling, networking, physical layout, and data-center or lab constraints.
- Contribute to platform design reviews, refresh cycles, technical documentation, and cross-functional software/hardware/system architecture decisions.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- Doctorate with 0 years, master’s degree with 3 years, or bachelor’s degree with 5 years of related work experience.
- Strong experience developing HPC or systems software on Linux.
- Proficiency in Java, C++, or other system-level or performance-oriented languages.
- Hands-on experience with parallel computing, including MPI, OpenMP, or multithreading.
- Understanding of CPUs, memory hierarchies, storage, Ethernet, InfiniBand, and other HPC hardware fundamentals.
- Practical experience with clusters, servers, or rack-scale systems in laboratory or production environments.
- Strong debugging skills across software, operating-system, and hardware boundaries.
- Preferred experience with GPU computing such as CUDA or ROCm.
- Preferred experience with Docker, Singularity/Apptainer, Kubernetes in HPC contexts, high-speed interconnects, storage architectures, performance benchmarking, rack integration, and high-reliability systems.
- Preferred familiarity with MTBF, MTBA, and failure modes in large compute installations.
Benefits
- Base pay is listed as $105,900.00–$155,300.00 annually, with possible performance incentives and additional benefits.
- Benefits may include medical, dental, vision, life, voluntary benefits, 401(k) matching, ESPP, student debt assistance, tuition reimbursement, career development, financial planning, wellness and EAP programs, paid time off, holidays, and family care and bonding leave.
- The primary location is KLA’s Ann Arbor, Michigan site in the United States.
About KLA
KLA builds process control, inspection and metrology systems used to manufacture semiconductor wafers, reticles, advanced packaging, printed circuit boards, and displays, and sells related software and services to chipmakers and electronics manufacturers. Its business model centers on capital equipment sales plus service and support contracts. Headquartered in Milpitas, California, KLA traces its origins to the 1997 KLA-Tencor merger and is publicly traded on NASDAQ under the ticker KLAC.
