
Member of Technical Staff (AI Inference Engineer)
Perplexity6 months ago
Palo Alto, CA, USA +2 moreMid Level
H1B sponsor
Base Salary
$220k - $485k/yr
Responsibilities
- Support transformer-based retrieval, text-generation, and multimodal models in the inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
- Migrate in-house CUDA kernels to NVIDIA CuTe DSL for current and future GPU platforms.
- Develop the internal Rust-based inference server to improve serving performance and scalability.
- Profile and resolve bottlenecks across network ingress, continuous batching, and GPU kernel interleaving.
- Build dashboards, alerts, and automated remediation for inference reliability and regression detection.
- Respond to production incidents and apply lessons learned to improve system reliability.
Requirements
- At least 3 years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Deep experience with GPU programming and performance optimization using CUDA, Triton, CUTLASS, or similar technologies.
- Understanding of modern LLM architectures and the ability to deploy them reliably in production.
- Experience building and operating production distributed systems under real load, preferably performance-critical systems.
- Ability to work across Rust, Python, CUDA, and CuTe DSL.
- Familiarity with at least one deep learning framework: PyTorch, JAX, or TensorFlow.
- Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
- Understanding of LLM inference optimization techniques such as quantization, speculative decoding, and prefill-decode disaggregation.
- Experience with ML compiler or framework internals, distributed GPU communication, low-precision inference, profiling and debugging tools, or Kubernetes is beneficial.
Tech Stack
Categories
About Perplexity
Perplexity builds an AI-powered answer engine for web and mobile and an enterprise product, Perplexity Computer, for AI-driven workflows across tools and apps. The company monetizes through individual subscriptions (Perplexity Pro) and enterprise plans. Founded in 2022 and headquartered in San Francisco, it operates as a privately held company.