4 days ago
Houston, TX, USA or San Francisco, CA, USAMid Level
H1B Sponsor
Responsibilities
- Optimize end-to-end GPU performance for real-time autonomous driving workloads, including camera and LiDAR processing and neural network inference.
- Develop and optimize parallel computing algorithms and GPU-accelerated components.
- Design and improve onboard GPU software architectures for perception, planning, and control modules.
- Profile and analyze bottlenecks in GPU computation, memory access, data movement, synchronization, and CPU–GPU interaction.
- Debug and optimize GPU-based software for latency, throughput, resource utilization, and runtime stability on embedded platforms.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- Strong knowledge of parallel computing, GPU architecture, memory hierarchy, and performance optimization.
- Experience profiling GPU applications with NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.
- Experience deploying or optimizing neural network inference workloads with PyTorch, ONNX, and TensorRT.
- Experience with real-time embedded systems and large sensor data streams from cameras, LiDAR, and radar.
- Strong proficiency in C/C++ and Python.
- Preferred experience with CUDA, OpenCL, or Vulkan GPU programming and optimization.
- Preferred experience with NVIDIA Jetson Thor, NVIDIA DRIVE Thor, or similar embedded GPU platforms.
- Preferred experience with model quantization, including FP8 and NVFP4.
- Preferred experience managing concurrent GPU workloads and resource isolation with NVIDIA Multi-Process Service (MPS), Multi-Instance GPU (MIG), or related technologies.
- Preferred experience with GPU-accelerated sensor data compression.
