tiiuae/falcon-180B-chat
Visit HuggingFace for more details.
mistralai/Mistral-7B-v0.3
apache-2.0Model Card for Mistral-7B-v0.3 The Mistral-7B-v0.3 Large Language Model (LLM) is a Mistral-7B-v0.2 with extended vocabulary. Mistral-7B-v0.3...
google/gemma-2-27b-it
gemmaVisit HuggingFace for more details.
NovaSky-AI/Sky-T1-32B-Preview
apache-2.0Model Details Model Description This is a 32B reasoning model trained from Qwen2.5-32B-Instruct with 17K data. The performance is on par wit...
distilbert/distilgpt2
apache-2.0DistilGPT2 DistilGPT2 (short for Distilled-GPT2) is an English-language model pre-trained with the supervision of the smallest version of Ge...
deepseek-ai/deepseek-coder-33b-instruct
other🏠Homepage | 🤖 Chat with DeepSeek Coder | Discord | Wechat(微信) 1. Introduction of Deepseek Coder Deepseek Coder is composed of a series of ...
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
mitDeepSeek-R1 Paper Link👁️ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...
bigscience/bloomz
bigscience-bloom-rail-1.0!xmtf( Table of Contents 1. Model Summary(model-summary) 2. Use(use) 3. Limitations(limitations) 4. Training(training) 5. Evaluation(evaluat...
NousResearch/Hermes-2-Pro-Mistral-7B
apache-2.0Hermes 2 Pro - Mistral 7B !image/png( Model Description Hermes 2 Pro on Mistral 7B is the new flagship 7B Hermes! Hermes 2 Pro is an upgrade...
Qwen/Qwen2.5-Coder-7B-Instruct
apache-2.0Qwen2.5-Coder-7B-Instruct Introduction Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as Cod...