HuggingFaceTB/SmolVLM-256M-Instruct
apache-2.0SmolVLM-256M SmolVLM-256M is the smallest multimodal model in the world. It accepts arbitrary sequences of image and text inputs to produce ...
CohereLabs/aya-vision-32b
cc-by-nc-4.0Visit HuggingFace for more details.
stepfun-ai/GOT-OCR-2.0-hf
apache-2.0General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers π€ implementation π€ Spaces Demo( | πGitHub( | πPaper...
moonshotai/Kimi-VL-A3B-Instruct
mitπ Tech Report | π Github | π¬ Chat Web Introduction We present Kimi-VL, an efficient open-source Mixture-of-Expert...
deepseek-ai/deepseek-vl2-tiny
other1. Introduction Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly imp...
OpenGVLab/InternVL3-78B
otherInternVL3-78B \π GitHub\( \π InternVL 1.0\( \π InternVL 1.5\( \π InternVL 2.5\( \π InternVL2.5-MPO\( \π InternVL3\( \π Blog\( \π¨οΈ Ch...
microsoft/OmniParser
mitπ’ Project Page( Blog Post( Demo( Model Summary OmniParser is a general screen parsing tool, which interprets/converts UI screenshot to stru...
ByteDance-Seed/UI-TARS-7B-SFT
apache-2.0UI-TARS-7B-SFT UI-TARS-2B-SFT( | UI-TARS-7B-SFT( | UI-TARS-7B-DPO( | UI-TARS-72B-SFT( | UI-T...
meta-llama/Llama-4-Scout-17B-16E
otherVisit HuggingFace for more details.
deepseek-ai/deepseek-vl2-small
other1. Introduction Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly imp...