Qwen/Qwen2.5-VL-72B-Instruct

other

--- license: other licensename: qwen licenselink: language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformer...

image text to textBy Qwen

OpenGVLab/InternVL-Chat-V1-5

mit

InternVL-Chat-V1-5 \📂 GitHub\( \📜 InternVL 1.0\( \📜 InternVL 1.5\( \📜 Mini-InternVL\( \📜 InternVL 2.5\( \🆕 Blog\( \🗨️ Chat Demo\( \🤗...

image text to textBy OpenGVLab

moonshotai/Kimi-VL-A3B-Thinking

mit

> !Warning > This model has a new version: Kimi-VL-A3B-Thinking-2506( Please consider using the new 2506 version for better abilties on gene...

image text to textBy moonshotai

microsoft/Magma-8B

mit

Model Card for Magma-8B Magma: A Foundation Model for Multimodal AI Agents Jianwei Yang( Reuben Tan( Qianhui Wu( Ruijie Zheng( Baolin Peng( ...

image text to textBy microsoft

Qwen/Qwen2.5-VL-3B-Instruct

--- licensename: qwen-research licenselink: language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers --- ...

image text to textBy Qwen

google/gemma-3-12b-it

gemma

Visit HuggingFace for more details.

image text to textBy google

Salesforce/blip2-opt-2.7b

mit

BLIP-2, OPT-2.7b, pre-trained only BLIP-2 model, leveraging OPT-2.7b( (a large language model with 2.7 billion parameters). It was introduce...

image text to textBy Salesforce

Qwen/Qwen2.5-VL-32B-Instruct

apache-2.0

Qwen2.5-VL-32B-Instruct Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and prob...

image text to textBy Qwen

google/gemma-3n-E4B-it-litert-preview

gemma

Visit HuggingFace for more details.

image text to textBy google

microsoft/Florence-2-large-ft

mit

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks Model Summary This Hub repository contains a HuggingFace's tran...

image text to textBy microsoft
Showing 10 of 159