meta-llama/Llama-4-Scout-17B-16E-Instruct
otherVisit HuggingFace for more details.
Qwen/Qwen2.5-VL-7B-Instruct
apache-2.0--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers --- Qwen2.5-VL-7B-Instru...
nvidia/NVLM-D-72B
cc-by-nc-4.0Model Overview Description This family of models performs vision-language and text-only tasks including optical character recognition, multi...
allenai/olmOCR-7B-0225-preview
apache-2.0olmOCR-7B-0225-preview This is a preview release of the olmOCR model that's fine tuned from Qwen2-VL-7B-Instruct using the olmOCR-mix-0225( ...
mistralai/Pixtral-12B-2409
apache-2.0Model Card for Pixtral-12B-2409 The Pixtral-12B-2409 is a Multimodal Model of 12B parameters plus a 400M parameter vision encoder. For more ...
rhymes-ai/Aria
apache-2.0Aria --> Aria Model Card Dec 1, 2024 We have released the base models (with native multimodal pre-training) for Aria (Aria-Base-8K( and Aria...
meta-llama/Llama-3.2-11B-Vision
llama3.2Visit HuggingFace for more details.
liuhaotian/llava-v1.5-13b
LLaVA Model Card Model details Model type: LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal i...
HuggingFaceTB/SmolVLM-Instruct
apache-2.0SmolVLM SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Design...
liuhaotian/llava-v1.5-7b
LLaVA Model Card Model details Model type: LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal i...