Anyscale

Distributed LLM Inference Engineer

Anyscale
Apply
3 months ago
Palo Alto, CA, USA or San Francisco, CA, USAMid Level
H1B Sponsor

Base Salary

$170k - $245k/yr

Responsibilities

  • Develop and rapidly ship end-to-end batch and online inference solutions for large-scale use by Ray users and Anyscale customers.
  • Integrate Ray Data with LLM engines and implement optimizations for high-throughput, low-latency, and cost-efficient ML inference.
  • Integrate open-source software such as vLLM, collaborate with its community, and contribute improvements upstream.
  • Track current open-source and research developments and implement or extend state-of-the-art inference practices.

Requirements

  • Familiarity with running large-scale ML inference with high throughput and low latency.
  • Familiarity with deep learning and frameworks such as PyTorch.
  • Solid understanding of distributed systems and ML inference challenges.
  • Preferred qualifications include ML systems knowledge, experience with Ray, work with vLLM or TensorRT-LLM, contributions to PyTorch or TensorFlow, contributions to Triton, TVM, or MLIR, and experience with GPUs or CUDA.

Benefits

  • Stock options and eligibility for Anyscale equity and benefits offerings.
  • Healthcare premiums covered by Anyscale at 99% for employees and dependents.
  • 401k retirement plan, education and wellbeing stipend, paid parental leave, fertility benefits, and paid time off.
  • Commute reimbursement and fully covered in-office meals.

Tech Stack

PyTorchTensorFlow
Anyscale

About Anyscale

201-500 employees

Anyscale enables Python developers to build and run all their AI—from data prep to training and inference—at any scale. Anyscale is trusted by leading AI teams at Canva, TripAdvisor, Physical Intelligence, Coinbase and more.

Contact me