HuggingFaceTB/SmolVLM-256M-Instruct

apache-2.0

SmolVLM-256M SmolVLM-256M is the smallest multimodal model in the world. It accepts arbitrary sequences of image and text inputs to produce ...

image text to textBy HuggingFaceTB

CohereLabs/aya-vision-32b

cc-by-nc-4.0

Visit HuggingFace for more details.

image text to textBy CohereLabs

stepfun-ai/GOT-OCR-2.0-hf

apache-2.0

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers πŸ€— implementation πŸ€— Spaces Demo( | 🌟GitHub( | πŸ“œPaper...

image text to textBy stepfun-ai

moonshotai/Kimi-VL-A3B-Instruct

mit

πŸ“„ Tech Report  |  πŸ“„ Github  |  πŸ’¬ Chat Web Introduction We present Kimi-VL, an efficient open-source Mixture-of-Expert...

image text to textBy moonshotai

deepseek-ai/deepseek-vl2-tiny

other

1. Introduction Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly imp...

image text to textBy deepseek-ai

OpenGVLab/InternVL3-78B

other

InternVL3-78B \πŸ“‚ GitHub\( \πŸ“œ InternVL 1.0\( \πŸ“œ InternVL 1.5\( \πŸ“œ InternVL 2.5\( \πŸ“œ InternVL2.5-MPO\( \πŸ“œ InternVL3\( \πŸ†• Blog\( \πŸ—¨οΈ Ch...

image text to textBy OpenGVLab

microsoft/OmniParser

mit

πŸ“’ Project Page( Blog Post( Demo( Model Summary OmniParser is a general screen parsing tool, which interprets/converts UI screenshot to stru...

image text to textBy microsoft

ByteDance-Seed/UI-TARS-7B-SFT

apache-2.0

UI-TARS-7B-SFT UI-TARS-2B-SFT(  |  UI-TARS-7B-SFT(  |  UI-TARS-7B-DPO(  |  UI-TARS-72B-SFT(  |  UI-T...

image text to textBy ByteDance-Seed

meta-llama/Llama-4-Scout-17B-16E

other

Visit HuggingFace for more details.

image text to textBy meta-llama

deepseek-ai/deepseek-vl2-small

other

1. Introduction Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly imp...

image text to textBy deepseek-ai
Showing 10 of 159