1 day ago
London, United KingdomSenior
Responsibilities
- Own model releases from initial requirements through training, evaluation, iteration, and deployment readiness.
- Train and iterate on PyTorch models using experiments, ablations, and defined evaluation criteria.
- Debug regressions, identify root causes, and improve model performance.
- Apply model optimization techniques such as quantization and distillation.
- Collaborate with ML and performance engineering teams on handoffs, bottlenecks, and optimization priorities.
- Communicate delivery timelines, trade-offs, and readiness criteria with stakeholders.
Requirements
- Proven experience improving production-system performance under constraints such as latency, memory, bandwidth, power, thermal limits, or cost.
- Strong hands-on experience training and iterating on deep learning models in PyTorch.
- Strong proficiency with at least one relevant toolchain, such as TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL.
- Ability to reason from high-level model behavior through low-level kernel and runtime execution.
- Familiarity with model optimization concepts including quantization and distillation.
- Strong engineering fundamentals and collaboration skills.
- Experience with edge, embedded, real-time, or similarly constrained ML production systems is desirable.
- Exposure to ML workflows spanning training, evaluation, and deployment handoff is desirable.
- Exposure to benchmarking ML models on real devices and managing system-level constraints is desirable.
