deepseek-ai/DeepSeek-Prover-V2-671B
1. Introduction We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with ini...
PygmalionAI/pygmalion-6b
creativeml-openrail-mPygmalion 6B Model description Pymalion 6B is a proof-of-concept dialogue model based on EleutherAI's GPT-J-6B( Warning: This model is NOT s...
Gustavosta/MagicPrompt-Stable-Diffusion
mitMagicPrompt - Stable Diffusion This is a model from the MagicPrompt series of models, which are GPT-2( models intended to generate prompt te...
deepseek-ai/DeepSeek-R1-Distill-Llama-8B
mitDeepSeek-R1 Paper Link👁️ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...
mistralai/Mixtral-8x22B-Instruct-v0.1
apache-2.0Model Card for Mixtral-8x22B-Instruct-v0.1 Encode and Decode with mistralcommon py from mistralcommon.tokens.tokenizers.mistral import Mistr...
Qwen/Qwen2-72B-Instruct
otherQwen2-72B-Instruct Introduction Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language model...
togethercomputer/GPT-NeoXT-Chat-Base-20B
apache-2.0Feel free to try out our OpenChatKit feedback app( GPT-NeoXT-Chat-Base-20B-v0.16 > TLDR: As part of OpenChatKit (codebase available here( > ...
Open-Orca/Mistral-7B-OpenOrca
apache-2.0🐋 Mistral-7B-OpenOrca 🐋 !OpenOrca Logo( "MistralOrca Logo") ( OpenOrca - Mistral - 7B - 8k We have used our own OpenOrca dataset( to fine-...
Qwen/Qwen2.5-7B-Instruct
apache-2.0Qwen2.5-7B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base langua...
deepseek-ai/DeepSeek-R1-Distill-Llama-70B
mitDeepSeek-R1 Paper Link👁️ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...