3 hours ago
Responsibilities
- Design and implement inference algorithms for LLMs, including sampling, speculative decoding, model compression, knowledge distillation, and quantization techniques.
- Develop and optimize low-level software and custom kernels to improve AI model latency, throughput, and memory usage across hardware and model architectures.
- Co-design algorithms with hardware architects to exploit custom hardware and accelerator capabilities.
- Lead a small technical team and collaborate closely with other Google teams.
- Publish and disseminate research to influence the broader AI research ecosystem.
Requirements
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience testing and launching software products and 3 years of experience with software design and architecture.
- 5 years of experience in speech/audio, reinforcement learning, ML infrastructure, or another ML specialization.
- 5 years of experience with ML design and ML infrastructure, including model deployment, evaluation, data processing, debugging, or fine-tuning.
- Experience integrating generative AI tools or LLM interfaces into workflows.
Categories
AI ResearchML Engineering
About Google
A problem isn't truly solved until it's solved for all. Googlers build products that help create opportunities for everyone, whether down the street or across the globe. Bring your insight, imagination and a healthy disregard for the impossible. Bring everything that makes you unique. Together, we can build for everyone. Check out our career opportunities at goo.gle/3DLEokh
