4 months ago
Responsibilities
- Run experiments self-hosting models across cloud instances and on-premises hardware configurations.
- Develop reproducible and measurable deployment patterns optimized for latency, throughput, and cost.
- Improve orchestration and routing software, including caching, request scheduling, batching, and resource allocation.
- Integrate accelerator support, inference engine features, and infrastructure upgrades while identifying bottlenecks and required capabilities.
- Build and maintain benchmarking harnesses, regression suites, and performance dashboards.
- Translate performance data into actionable feedback and guide production-readiness and stack-optimization decisions.
Requirements
- Experience deploying and benchmarking large-model inference in production or production-equivalent environments.
- Familiarity with multi-node GPU deployments and associated networking and communication stacks.
- Strong end-to-end performance characterization skills, including isolating bottlenecks in networks, runtimes, memory subsystems, or models.
- Familiarity with serving frameworks such as Dynamo, Triton Inference Server, or similar orchestration layers.
- Clear communication skills and the ability to turn performance data into prioritized feedback.
- A disciplined and systematic approach to reproducible deployment, measurement methodology, and controlled comparisons.
Benefits
- Competitive salary determined by skills and experience.
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the London office with provided tools, workspace, and setup.
