meta-llama/Llama-3.2-3B-Instruct

llama3.2

Visit HuggingFace for more details.

text generationBy meta-llama

deepseek-ai/DeepSeek-R1-Distill-Qwen-32B

mit

DeepSeek-R1 Paper Link👁️ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...

text generationBy deepseek-ai

microsoft/phi-1_5

mit

Model Summary The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1( ...

text generationBy microsoft

01-ai/Yi-34B

apache-2.0

Building the Next Generation of Open-Source and Bilingual LLMs 🤗 Hugging Face • 🤖 ModelScope • ✡️ WiseModel 👩‍🚀 Ask questions or discuss...

text generationBy 01-ai

mistralai/Codestral-22B-v0.1

other

Model Card for Codestral-22B-v0.1 Encode and Decode with mistralcommon py from mistralcommon.tokens.tokenizers.mistral import MistralTokeniz...

text generationBy mistralai

cognitivecomputations/dolphin-2.5-mixtral-8x7b

apache-2.0

Dolphin 2.5 Mixtral 8x7b 🐬 !Discord( Discord: This model's training was sponsored by convai( This model is based on Mixtral-8x7b The base m...

text generationBy cognitivecomputations

deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

mit

DeepSeek-R1 Paper Link👁️ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...

text generationBy deepseek-ai

microsoft/Phi-3-mini-4k-instruct

mit

🎉 Phi-3.5: mini-instruct( MoE-instruct( ; vision-instruct( Model Summary The Phi-3-Mini-4K-Instruct is a 3.8B parameters, lightweight, stat...

text generationBy microsoft

ai21labs/Jamba-v0.1

apache-2.0

This is the base version of the Jamba model. We’ve since released a better, instruct-tuned version, Jamba-1.5-Mini( For even greater perform...

text generationBy ai21labs

tiiuae/falcon-180B

Visit HuggingFace for more details.

text generationBy tiiuae
Showing 10 of 563