3 months ago
Palo Alto, CA, USASenior
Base Salary
$185k - $300k/yr
Responsibilities
- Design and optimize inference pipelines and implement advanced acceleration techniques for efficient model serving
- Engineer GPU strategies across tensor, sequence, and pipeline parallelism to maximize efficiency and scalability
- Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL
- Bring video-generation and large language models into production with research and engineering partners
- Contribute to model training speed, stability, and resource-utilization improvements
- Lead code reviews, participate in technical discussions, and mentor engineers on inference and GPU programming best practices
Requirements
- 5+ years of engineering experience with inference acceleration and model deployment at scale
- Expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks
- Deep knowledge of GPU programming with CUDA and NCCL and experience with sequence, tensor, and pipeline parallelism
- Familiarity with video-generation models and large language models
- Experience collaborating across research and engineering teams and driving shared technical goals
- Experience with high-throughput video or real-time streaming model deployment is preferred
- Familiarity with distributed training and optimization toolkits is preferred
- Contributions to open-source AI infrastructure or deep learning compilers are preferred
- Startup or rapid prototyping experience is preferred
Benefits
- Equity in a fast-growing startup
- Comprehensive health benefits and monthly stipends
- Company retreats
- Palo Alto office-based work 3–5 days per week
Categories
About Pika
Pika builds an AI-powered idea-to-video platform that generates and edits short videos from text, images, or rough clips for creators and teams. The company develops its own generative video models and offers a web app used for storytelling, social content, and design workflows. Founded in 2023 and based in Palo Alto, it is a privately held startup in the generative video space.
