Qwen/Qwen2.5-VL-7B-Instruct
apache-2.0--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers --- Qwen2.5-VL-7B-Instru...
Florence-2-base
mitOfficial image-text-to-text model by microsoft.
nvidia/NVLM-D-72B
cc-by-nc-4.0Model Overview Description This family of models performs vision-language and text-only tasks including optical character recognition, multi...
allenai/olmOCR-7B-0225-preview
apache-2.0olmOCR-7B-0225-preview This is a preview release of the olmOCR model that's fine tuned from Qwen2-VL-7B-Instruct using the olmOCR-mix-0225( ...
mistralai/Pixtral-12B-2409
apache-2.0Model Card for Pixtral-12B-2409 The Pixtral-12B-2409 is a Multimodal Model of 12B parameters plus a 400M parameter vision encoder. For more ...
rhymes-ai/Aria
apache-2.0Aria --> Aria Model Card Dec 1, 2024 We have released the base models (with native multimodal pre-training) for Aria (Aria-Base-8K( and Aria...
microsoft/Florence-2-large
mitFlorence-2: Advancing a Unified Representation for a Variety of Vision Tasks Model Summary This is a continued pretrained version of Florenc...
stepfun-ai/GOT-OCR2_0
apache-2.0General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model 🔋Online Demo( | 🌟GitHub( | 📜Paper( Haoran Wei( Chenglong Liu, Jinyue C...
meta-llama/Llama-3.2-11B-Vision-Instruct
llama3.2Visit HuggingFace for more details.
ds4sd/SmolDocling-256M-preview
cdla-permissive-2.0📢 New Release: We’ve released granite-docling-258M, the successor to SmolDocling. It will now receive updates and support, check it out! Sm...