Qwen/Qwen2-72B-Instruct
otherQwen2-72B-Instruct Introduction Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language model...
togethercomputer/GPT-NeoXT-Chat-Base-20B
apache-2.0Feel free to try out our OpenChatKit feedback app( GPT-NeoXT-Chat-Base-20B-v0.16 > TLDR: As part of OpenChatKit (codebase available here( > ...
MiniMaxAI/MiniMax-Text-01
WeChat MiniMax-Text-01 1. Introduction MiniMax-Text-01 is a powerful language model with 456 billion total parameters, of which 45.9 billion...
Qwen/Qwen2.5-7B-Instruct
apache-2.0Qwen2.5-7B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base langua...
deepseek-ai/DeepSeek-R1-Distill-Llama-70B
mitDeepSeek-R1 Paper Link๐๏ธ 1. Introduction We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-...
google/gemma-2-9b
gemmaVisit HuggingFace for more details.
agentica-org/DeepCoder-14B-Preview
mitDeepCoder-14B-Preview ๐ Democratizing Reinforcement Learning for LLMs (RLLM) ๐ DeepCoder Overview DeepCoder-14B-Preview is a code reasonin...
allenai/OLMo-7B
apache-2.0!mof-class1-qualified( Model Card for OLMo 7B For transformers versions v4.40.0 or newer, we suggest using OLMo 7B HF( instead. OLMo is a se...
jinaai/ReaderLM-v2
cc-by-nc-4.0Trained by Jina AI. Blog( | API( | Colab( | AWS( | Azure( Arxiv( ReaderLM-v2 ReaderLM-v2 is a 1.5B parameter language model that converts ra...
Qwen/Qwen2-7B-Instruct
apache-2.0Qwen2-7B-Instruct Introduction Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models...