mistralai/Mixtral-8x7B-Instruct-v0.1
apache-2.0Model Card for Mixtral-8x7B Tokenization with mistral-common py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mi...
meta-llama/Llama-2-7b
llama2Visit HuggingFace for more details.
meta-llama/Llama-3.1-8B-Instruct
llama3.1Visit HuggingFace for more details.
meta-llama/Meta-Llama-3-8B-Instruct
llama3Visit HuggingFace for more details.
deepseek-ai/DeepSeek-V3
Paper Link👁️ 1. Introduction We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B a...
mistralai/Mistral-7B-v0.1
apache-2.0Model Card for Mistral-7B-v0.1 The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameter...
tiiuae/falcon-40b
apache-2.0🚀 Falcon-40B Falcon-40B is a 40B parameters causal decoder-only model built by TII( and trained on 1,000B tokens of RefinedWeb( enhanced wi...
microsoft/phi-2
mitModel Summary Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5( augmented with a ne...
mistralai/Mixtral-8x7B-v0.1
apache-2.0Model Card for Mixtral-8x7B The Mixtral-8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts. The Mistral-8x...
google/gemma-7b
gemmaVisit HuggingFace for more details.