Perplexity

Member of Technical Staff (AI Inference Engineer)

Perplexity
Apply
4 months ago
Palo Alto, CA, USA +2 moreMid Level
H1B Sponsor

Base Salary

$220k - $485k/yr

Responsibilities

  • Support transformer-based retrieval, text-generation, and multimodal models in the inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
  • Migrate in-house CUDA kernels to NVIDIA CuTe DSL for current and future GPU platforms.
  • Develop the internal Rust-based inference server to improve serving performance and scalability.
  • Profile and resolve bottlenecks across network ingress, continuous batching, and GPU kernel interleaving.
  • Build dashboards, alerts, and automated remediation for inference reliability and regression detection.
  • Respond to production incidents and apply lessons learned to improve system reliability.

Requirements

  • At least 3 years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Deep experience with GPU programming and performance optimization using CUDA, Triton, CUTLASS, or similar technologies.
  • Understanding of modern LLM architectures and the ability to deploy them reliably in production.
  • Experience building and operating production distributed systems under real load, preferably performance-critical systems.
  • Ability to work across Rust, Python, CUDA, and CuTe DSL.
  • Familiarity with at least one deep learning framework: PyTorch, JAX, or TensorFlow.
  • Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
  • Understanding of LLM inference optimization techniques such as quantization, speculative decoding, and prefill-decode disaggregation.
  • Experience with ML compiler or framework internals, distributed GPU communication, low-precision inference, profiling and debugging tools, or Kubernetes is beneficial.
Perplexity

About Perplexity

201-500 employees

The most powerful answer engine. Powering curiosity with answers backed by up-to-date sources. This is where knowledge begins.