1 day ago
London, United KingdomStaff+
Responsibilities
- Profile the full inference stack to identify bottlenecks in model graphs, compilers, runtimes, kernel execution, and memory movement.
- Implement and validate compiler, runtime, and kernel optimizations, including operator fusion, scheduling, quantization-aware performance improvements, and custom kernels.
- Build benchmarking and regression testing to validate performance across models, devices, and software releases.
- Optimize ML inference for NVIDIA Orin/Thor, Qualcomm, edge accelerators, GPUs, and in-vehicle compute.
- Collaborate with model developers on architecture and training/deployment decisions affecting on-device performance.
- Contribute to technical roadmaps and tooling, raise performance engineering standards, and help provide technical direction.
Requirements
- Demonstrated experience improving production-system performance under tight latency, memory, bandwidth, power, thermal, or cost constraints.
- Strong proficiency with at least one relevant toolchain such as TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL, with the ability to learn adjacent frameworks.
- Ability to work across high-level model behavior and low-level kernel/runtime execution.
- Strong software engineering fundamentals in debugging, profiling, testing, and maintainable code.
- Clear communication and collaborative stakeholder management around performance trade-offs and priorities.
- Exposure to embedded or edge ML deployment, benchmarking on real devices, and system-level constraints is desirable.
- Experience with NVIDIA and/or Qualcomm SoCs and performance tooling is desirable.
- Python and C++ proficiency is desirable.
- Experience mentoring others or driving technical direction in a small, fast-moving team is desirable.
Benefits
- Full-time position based in Wayve’s London office.
- Hybrid working policy combining office and workshop collaboration with working from home.
- Inclusive interview experience with accommodations and adjustments available upon request.
