4 months ago
Responsibilities
- Contribute upstream to SGLang and vLLM while maintaining internal forks where needed.
- Improve inference-engine hardware awareness across scheduling, memory management, and execution.
- Design and implement parallelism and disaggregation strategies for heterogeneous hardware.
- Collaborate with accelerator systems software engineers to align engine abstractions with diverse hardware capabilities.
Requirements
- Deep familiarity with SGLang, vLLM, or comparable inference-serving frameworks, including scheduler design, memory management, and execution pipelines.
- Strong experience with high-performance Python and C++/CUDA systems in the context of machine-learning inference.
- Experience designing or implementing parallelism strategies for large-model serving.
- Understanding of disaggregated serving architectures and their tradeoffs.
- A demonstrated ability to work effectively in fast-moving open-source codebases with evolving APIs and design conventions.
Benefits
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the company’s London office with provided tools, space, and equipment.
