1 day ago
London, United KingdomStaff+
Responsibilities
- Identify, implement, and validate optimizations in ML compilers, runtimes, and kernels, including operator fusion, scheduling, quantization-aware performance, and custom kernels.
- Profile the full inference stack to locate bottlenecks and deliver measurable improvements in latency, memory, bandwidth, power, thermal performance, or cost.
- Build benchmarking and regression-testing systems across models, devices, and software releases.
- Develop and optimize ML inference for NVIDIA Orin/Thor, Qualcomm, and other target platforms.
- Collaborate with model developers on architecture and training/deployment decisions affecting on-device performance.
- Contribute to technical roadmaps and tooling and raise performance-engineering standards across the team.
- Collaborate with cross-functional stakeholders and potentially mentor others while driving technical direction.
Requirements
- Proven experience improving performance in production systems subject to tight resource or performance constraints.
- Strong proficiency with at least one relevant stack or toolchain, such as TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, or ONNX.
- Ability to work across abstraction levels from high-level model behavior to low-level kernel and runtime execution.
- Strong software engineering fundamentals, including debugging, profiling, testing, and maintainable code.
- Clear communication and collaboration skills for aligning stakeholders on performance trade-offs and priorities.
- Experience with compute graph scheduling and execution across multiple targets is desirable.
- Experience deploying ML models to embedded or edge devices, benchmarking on real devices, and handling system-level constraints is desirable.
- Experience with NVIDIA or Qualcomm SoCs and performance tooling is desirable.
- Python and C++ proficiency is desirable.
- Experience mentoring others or driving technical direction in a small, fast-moving team is desirable.
