microsoft/biogpt
mitBioGPT Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the gene...
microsoft/Phi-4-reasoning-plus
mitPhi-4-reasoning-plus Model Card Phi-4-reasoning Technical Report( Model Summary | | | |-------------------------|---------------------------...
cerebras/btlm-3b-8k-base
apache-2.0Visit HuggingFace for more details.
chavinlo/alpaca-native
Stanford Alpaca This is a replica of Alpaca by Stanford' tatsu Trained using the original instructions with a minor modification in FSDP mod...
KoboldAI/OPT-13B-Erebus
otherOPT 13B - Erebus Model description This is the second generation of the original Shinen made by Mr. Seeker. The full dataset consists of 6 d...
Qwen/Qwen2.5-3B-Instruct
otherQwen2.5-3B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base langua...
openai-community/openai-gpt
mitOpenAI GPT 1 Table of Contents - Model Details(model-details) - How To Get Started With the Model(how-to-get-started-with-the-model) - Uses(...
TheBloke/Mistral-7B-v0.1-GGUF
apache-2.0Chat & support: TheBloke's Discord server Want to contribute? TheBloke's Patreon page TheBloke's LLM work is generously supported by a grant...
CohereLabs/aya-expanse-32b
cc-by-nc-4.0Visit HuggingFace for more details.
XiaomiMiMo/MiMo-7B-RL
mit━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language ModelFrom Pretraining to Posttraining ━━━━━━━━━━━━━━...