1bitLLM/bitnet_b1_58-3B
mitThis is a reproduction of the BitNet b1.58 paper. The models are trained with RedPajama dataset for 100B tokens. The hypers, as well as two-...
THUDM/codegeex4-all-9b
otherCodeGeeX4: Open Multilingual Code Generation Model δΈζ(./READMEzh.md) GitHub( We introduce CodeGeeX4-ALL-9B, the open-source version of the l...
nvidia/DeepSeek-R1-FP4
mitModel Overview Description: The NVIDIA DeepSeek R1 FP4 model is the quantized version of the DeepSeek AI's DeepSeek R1 model, which is an au...
Qwen/Qwen-VL
Qwen-VL Qwen-VL π€ π€  ο½ Qwen-VL-Chat π€ π€  (Int4: π€ π€ ) ο½ Qwen-VL-Plus π€ π€  ο½ Qwen-VL-Max π€ π€  Web   | &...
lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF
llama3.1π« Community Model> Llama 3.1 8B Instruct by Meta πΎ LM Studio( Community models highlights program. Highlighting new & noteworthy models by...
alpindale/goliath-120b
llama2Goliath 120B An auto-regressive causal LM created by combining 2x finetuned Llama-2 70B( into one. Please check out the quantized formats pr...
nisten/Biggie-SmoLlm-0.15B-Base
mitTINY Frankenstein of SmolLM-135M( upped to 0.18b Use this frankenbase for training. Sorry for the mislabelling, the model is a 0.18b 181m pa...
Qwen/Qwen2.5-14B-Instruct
apache-2.0Qwen2.5-14B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base langu...
meta-llama/LlamaGuard-7b
llama2Visit HuggingFace for more details.
tiiuae/falcon-mamba-7b
otherTable of Contents 0. TL;DR(TL;DR) 1. Model Details(model-details) 2. Usage(usage) 3. Training Details(training-details) 4. Evaluation(evalua...