6 hours ago
London, United KingdomSenior
Responsibilities
- Collaborate with research teams to ensure next-generation modeling approaches are designed and implemented with production considerations.
- Work with infrastructure teams to deliver efficient, high-performance serving infrastructure and address bottlenecks in speed, scale, and quality.
- Automate tasks, eliminate redundancies, build performant tests, and improve the velocity of model releases.
- Develop deep expertise in serving frameworks, preprocessing pipelines, caching mechanisms, and related technologies.
- Use roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and GPU/TPU serving infrastructure.
- Optimize and deploy large language models such as Gemini onto production infrastructure.
Requirements
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 2 years of experience deploying and maintaining machine learning models in a live production environment.
- Experience profiling, configuring, or executing ML workloads directly on hardware accelerators such as GPUs or TPUs.
- Experience designing, building, or optimizing model-serving infrastructure or inference backends.
- Preferred experience developing serving infrastructure.
- Preferred experience programming GPUs or TPUs through JAX, PyTorch, Pallas, CUDA, OpenCL, or similar models.
- Preferred experience profiling software to identify performance bottlenecks.
- Preferred experience optimizing distributed ML systems and parallelism, including data, model, or pipeline parallelism.
- Familiarity with writing performance-optimized kernels.
- Understanding of LLM architecture and inference performance dynamics, including Transformer models, memory bandwidth, compute bounds, and KV-cache scaling.
Benefits
- Opportunities for both individual contributor and technical leadership work.
- Open to applicants from software engineering and research engineering backgrounds, including specialist and generalist serving interests.
- Work with interdisciplinary research and engineering teams on AI systems intended for broad public benefit and scientific discovery.
Tech Stack
Categories
About Google
Google builds consumer and enterprise software and services including Search, Android, YouTube, Chrome, Maps, Gmail, and Google Cloud. Its business model centers on digital advertising and paid cloud, software, and hardware offerings (e.g., Pixel and Nest) for consumers, developers, and organizations. Founded in 1998 and headquartered in Mountain View, California, Google operates globally as a subsidiary of Alphabet Inc.
