deepseek-ai/Janus-Pro-7B

mit

1. Introduction Janus-Pro is a novel autoregressive framework that unifies multimodal understanding and generation. It addresses the limitat...

any to anyBy deepseek-ai

Qwen/Qwen2.5-Omni-7B

other

Qwen2.5-Omni Overview Introduction Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, i...

any to anyBy Qwen

openbmb/MiniCPM-o-2_6

A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your Phone GitHub( | Online Demo( | Technical Blog( | Join Us( News ...

any to anyBy openbmb

deepseek-ai/Janus-Pro-1B

mit

1. Introduction Janus-Pro is a novel autoregressive framework that unifies multimodal understanding and generation. It addresses the limitat...

any to anyBy deepseek-ai

ByteDance-Seed/BAGEL-7B-MoT

apache-2.0

🥯 BAGEL • Unified Model for Multimodal Understanding and Generation > We present BAGEL, an open‑source multimodal foundation model with 7B ...

any to anyBy ByteDance-Seed

Qwen/Qwen2.5-Omni-3B

other

Qwen2.5-Omni Overview Introduction Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, i...

any to anyBy Qwen

BAAI/Emu3-Gen

apache-2.0

Emu3: Next-Token Prediction is All You Need Emu3 Team, BAAI( | Project Page( | Paper( | 🤗HF Models( | github( | Demo( | We introduce Emu3, ...

any to anyBy BAAI

deepseek-ai/Janus-Pro-7B

mit

1. Introduction Janus-Pro is a novel autoregressive framework that unifies multimodal understanding and generation. It addresses the limitat...

any to anyBy deepseek-ai

Qwen/Qwen2.5-Omni-7B

other

Qwen2.5-Omni Overview Introduction Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, i...

any to anyBy Qwen

openbmb/MiniCPM-o-2_6

A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your Phone GitHub( | Online Demo( | Technical Blog( | Join Us( News ...

any to anyBy openbmb
Showing 10 of 27