meta-llama/Llama-3.2-90B-Vision-Instruct

llama3.2

Visit HuggingFace for more details.

image text to textBy meta-llama

deepseek-ai/deepseek-vl2

other

1. Introduction Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly imp...

image text to textBy deepseek-ai

meta-llama/Llama-4-Maverick-17B-128E-Instruct

other

Visit HuggingFace for more details.

image text to textBy meta-llama

google/paligemma-3b-pt-224

gemma

Visit HuggingFace for more details.

image text to textBy google

Qwen/Qwen2-VL-72B-Instruct

other

Qwen2-VL-72B-Instruct Introduction We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year o...

image text to textBy Qwen

CohereLabs/aya-vision-8b

cc-by-nc-4.0

Visit HuggingFace for more details.

image text to textBy CohereLabs

google/gemma-3-27b-it-qat-q4_0-gguf

gemma

Visit HuggingFace for more details.

image text to textBy google

ByteDance-Seed/UI-TARS-1.5-7B

apache-2.0

--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal - gui libraryname: transformers --- UI-TARS-1.5 Mo...

image text to textBy ByteDance-Seed

llava-hf/llava-1.5-7b-hf

llama2

LLaVA Model Card !image/png( Below is the model card of Llava model 7b, which is copied from the original Llava model card that you can find...

image text to textBy llava-hf

MiniMaxAI/MiniMax-VL-01

WeChat MiniMax-VL-01 1. Introduction We are delighted to introduce our MiniMax-VL-01 model. It adopts the "ViT-MLP-LLM" framework, which is ...

image text to textBy MiniMaxAI
Showing 10 of 159