
Senior Staff Performance Engineer
Samsung Semiconductor23 hours ago
San Jose, CA, USAStaff+
Base Salary
$189k - $301k/yr
Responsibilities
- Build and operate AI environments representing production workloads, including agentic workflows, distributed inference, disaggregated serving architectures, and Mixture-of-Experts deployments.
- Collect workload traces, runtime telemetry, and performance data across the software stack.
- Characterize workloads and identify compute, memory, communication, and scheduling bottlenecks across applications, frameworks, runtimes, hosts, and devices.
- Communicate performance findings to hardware architects, systems engineers, and software researchers through reports, presentations, and architecture reviews.
- Define performance evaluation methodologies and benchmarking standards and set technical direction for workload characterization.
Requirements
- A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent industry experience.
- At least 10 years of relevant industry experience with a BS, 8 years with an MS, or 5 years with a PhD in performance engineering, AI systems, distributed systems, high-performance computing, or a related area.
- Additional experience and a record of technical leadership across teams may qualify candidates for Senior Staff consideration.
- Ability to interpret workload traces, runtime telemetry, and performance data, identify bottlenecks, and explain their causes across software and hardware layers.
- Knowledge of LLM serving and scheduling, attention and KV-cache management, kernel launch and memory-transfer overhead, collective communication, accelerator throughput, memory hierarchy, memory bandwidth, and interconnect topology.
- Experience characterizing agentic workflows, long-context processing, Mixture-of-Experts models, or disaggregated inference deployments.
- Experience profiling and optimizing AI workloads on NVIDIA GPU platforms using Nsight Systems and Nsight Compute.
- Experience analyzing multi-node AI deployments, including synchronization overhead, load imbalance, communication patterns, and scaling behavior.
- Experience with AI frameworks or serving systems such as PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, or Megatron-LM.
Benefits
- Base pay is listed separately in the posting, with incentive opportunities tied to individual and company performance.
- Medical, dental, vision, and 401(k) benefits are provided.
- At least four weeks of paid time off annually, plus holidays and sick leave.
- Charitable giving match and community involvement opportunities.
- Fertility-care or adoption stipend, caregiver support, and medical travel support.
- Emotional wellness support, including on-demand apps and confidential therapy sessions.
- Onsite café, gym, and virtual fitness classes.
- Flexible work environment; the role requires daily onsite presence at the San Jose, California office under the company’s Flexible Work policy.