Where can I deploy this model for inference?

by catworld1212 - opened Apr 27

Apr 27

Hi, I'm impressed with the work on InternVL and I'm interested in deploying its inference as an endpoint. Unfortunately, vLLM and TGI don't support this. Could anyone offer guidance on how to achieve this? I'd appreciate any suggestions you may have.

whai362

Apr 28

See "Chat Web Demo" at https://github.com/OpenGVLab/InternVL/blob/main/README.md

catworld1212

Apr 30

See "Chat Web Demo" at https://github.com/OpenGVLab/InternVL/blob/main/README.md

I want to deploy it as an inference not run it as a demo, Can you tell do InternVL-Chat-V1-5 requires flash attention?

catworld1212

May 2

Hi @whai362 @czczup what's the proper way to few-shot prompting (also called in-context learning? How do I give the previous context? I'm using lmdeploy to serve the inference can you help me, please?

czczup

OpenGVLab org Aug 22

Hi @whai362 @czczup what's the proper way to few-shot prompting (also called in-context learning? How do I give the previous context? I'm using lmdeploy to serve the inference can you help me, please?

Hi, you can perform few-shot prompting in the form of multi-turn dialogue, placing the few-shot examples in the conversation history. I think this shouldn't be different from how other VLMs perform few-shot prompting.

czczup changed discussion status to closed Aug 22

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment