add VLLM support? #1010

cboettig · 2024-09-22T04:38:50Z

Problem/Solution

It has been great to see Ollama added as a first-class option in #646, this has made it easy to access a huge variety of models and has been working very well us.

I increasingly see groups and university providers using VLLM for this as well. I'm out of my depth but I understand VLLM is considered better suited when a group is serving a local model to multiple users (e.g. from a local GPU cluster, rather than everyone running an independent Ollama). It gets passing mention in some threads here as well. I think supporting more providers is all to the good and would love to see support for this as a backend similar to the existing Ollama support, though maybe I'm not understanding the details and that is unnecessary? (i.e. it looks like it might be possible to simply use the OpenAI configuration with alternative endpoint to access a VLLM server?)

cboettig · 2024-10-03T03:48:06Z

It looks like the team at the National Research Platform has a nice work-around for this at the moment using LiteLLM via it's OpenAI-compatible API (https://docs.litellm.ai/docs/proxy/user_keys) This works, though isn't really the same as direct VLLM support, but thought it was worth mentioning.

cboettig added the enhancement New feature or request label Sep 22, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

add VLLM support? #1010

add VLLM support? #1010

cboettig commented Sep 22, 2024

cboettig commented Oct 3, 2024 •

edited

Loading

add VLLM support? #1010

add VLLM support? #1010

Comments

cboettig commented Sep 22, 2024

Problem/Solution

cboettig commented Oct 3, 2024 • edited Loading

cboettig commented Oct 3, 2024 •

edited

Loading