Qwen/Qwen2.5-VL-7B-Instruct

apache-2.0

--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers --- Qwen2.5-VL-7B-Instru...

image text to textBy Qwen

Florence-2-base

mit

Official image-text-to-text model by microsoft.

image text to textBy microsoft

nvidia/NVLM-D-72B

cc-by-nc-4.0

Model Overview Description This family of models performs vision-language and text-only tasks including optical character recognition, multi...

image text to textBy nvidia

allenai/olmOCR-7B-0225-preview

apache-2.0

olmOCR-7B-0225-preview This is a preview release of the olmOCR model that's fine tuned from Qwen2-VL-7B-Instruct using the olmOCR-mix-0225( ...

image text to textBy allenai

mistralai/Pixtral-12B-2409

apache-2.0

Model Card for Pixtral-12B-2409 The Pixtral-12B-2409 is a Multimodal Model of 12B parameters plus a 400M parameter vision encoder. For more ...

image text to textBy mistralai

rhymes-ai/Aria

apache-2.0

Aria --> Aria Model Card Dec 1, 2024 We have released the base models (with native multimodal pre-training) for Aria (Aria-Base-8K( and Aria...

image text to textBy rhymes-ai

microsoft/Florence-2-large

mit

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks Model Summary This is a continued pretrained version of Florenc...

image text to textBy microsoft

stepfun-ai/GOT-OCR2_0

apache-2.0

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model 🔋Online Demo( | 🌟GitHub( | 📜Paper( Haoran Wei( Chenglong Liu, Jinyue C...

image text to textBy stepfun-ai

meta-llama/Llama-3.2-11B-Vision-Instruct

llama3.2

Visit HuggingFace for more details.

image text to textBy meta-llama

ds4sd/SmolDocling-256M-preview

cdla-permissive-2.0

📢 New Release: We’ve released granite-docling-258M, the successor to SmolDocling. It will now receive updates and support, check it out! Sm...

image text to textBy ds4sd
Showing 10 of 159