meta-llama/Llama-3.1-8B
llama3.1Visit HuggingFace for more details.
meta-llama/Llama-3.2-3B-Instruct
llama3.2Visit HuggingFace for more details.
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
mitDeepSeek-R1 Paper LinkποΈ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...
microsoft/phi-1_5
mitModel Summary The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1( ...
01-ai/Yi-34B
apache-2.0Building the Next Generation of Open-Source and Bilingual LLMs π€ Hugging Face β’ π€ ModelScope β’ β‘οΈ WiseModel π©βπ Ask questions or discuss...
TinyLlama/TinyLlama-1.1B-Chat-v1.0
apache-2.0TinyLlama-1.1B The TinyLlama project aims to pretrain a 1.1B Llama model on 3 trillion tokens. With some proper optimization, we can achieve...
mistralai/Codestral-22B-v0.1
otherModel Card for Codestral-22B-v0.1 Encode and Decode with mistralcommon py from mistralcommon.tokens.tokenizers.mistral import MistralTokeniz...
cognitivecomputations/dolphin-2.5-mixtral-8x7b
apache-2.0Dolphin 2.5 Mixtral 8x7b π¬ !Discord( Discord: This model's training was sponsored by convai( This model is based on Mixtral-8x7b The base m...
microsoft/Phi-3-mini-4k-instruct
mitπ Phi-3.5: mini-instruct( MoE-instruct( ; vision-instruct( Model Summary The Phi-3-Mini-4K-Instruct is a 3.8B parameters, lightweight, stat...
ai21labs/Jamba-v0.1
apache-2.0This is the base version of the Jamba model. Weβve since released a better, instruct-tuned version, Jamba-1.5-Mini( For even greater perform...