
Member of Technical Staff (AI Inference Engineer)
Perplexity4 months ago
Palo Alto, CA, USA +2 moreMid Level
H1B Sponsor
Base Salary
$220k - $485k/yr
Responsibilities
- Support transformer-based retrieval, text-generation, and multimodal models in the inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
- Migrate in-house CUDA kernels to NVIDIA CuTe DSL for current and future GPU platforms.
- Develop the internal Rust-based inference server to improve serving performance and scalability.
- Profile and resolve bottlenecks across network ingress, continuous batching, and GPU kernel interleaving.
- Build dashboards, alerts, and automated remediation for inference reliability and regression detection.
- Respond to production incidents and apply lessons learned to improve system reliability.
Requirements
- At least 3 years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Deep experience with GPU programming and performance optimization using CUDA, Triton, CUTLASS, or similar technologies.
- Understanding of modern LLM architectures and the ability to deploy them reliably in production.
- Experience building and operating production distributed systems under real load, preferably performance-critical systems.
- Ability to work across Rust, Python, CUDA, and CuTe DSL.
- Familiarity with at least one deep learning framework: PyTorch, JAX, or TensorFlow.
- Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
- Understanding of LLM inference optimization techniques such as quantization, speculative decoding, and prefill-decode disaggregation.
- Experience with ML compiler or framework internals, distributed GPU communication, low-precision inference, profiling and debugging tools, or Kubernetes is beneficial.
Tech Stack
Categories
About Perplexity
The most powerful answer engine. Powering curiosity with answers backed by up-to-date sources. This is where knowledge begins.