1bitLLM/bitnet_b1_58-3B

mit

This is a reproduction of the BitNet b1.58 paper. The models are trained with RedPajama dataset for 100B tokens. The hypers, as well as two-...

text generationBy 1bitLLM

THUDM/codegeex4-all-9b

other

CodeGeeX4: Open Multilingual Code Generation Model δΈ­ζ–‡(./READMEzh.md) GitHub( We introduce CodeGeeX4-ALL-9B, the open-source version of the l...

text generationBy THUDM

nvidia/DeepSeek-R1-FP4

mit

Model Overview Description: The NVIDIA DeepSeek R1 FP4 model is the quantized version of the DeepSeek AI's DeepSeek R1 model, which is an au...

text generationBy nvidia

Qwen/Qwen-VL

Qwen-VL Qwen-VL πŸ€— πŸ€–&nbsp | Qwen-VL-Chat πŸ€— πŸ€–&nbsp (Int4: πŸ€— πŸ€–&nbsp) | Qwen-VL-Plus πŸ€— πŸ€–&nbsp | Qwen-VL-Max πŸ€— πŸ€–&nbsp Web&nbsp&nbsp | &...

text generationBy Qwen

lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF

llama3.1

πŸ’« Community Model> Llama 3.1 8B Instruct by Meta πŸ‘Ύ LM Studio( Community models highlights program. Highlighting new & noteworthy models by...

text generationBy lmstudio-community

alpindale/goliath-120b

llama2

Goliath 120B An auto-regressive causal LM created by combining 2x finetuned Llama-2 70B( into one. Please check out the quantized formats pr...

text generationBy alpindale

nisten/Biggie-SmoLlm-0.15B-Base

mit

TINY Frankenstein of SmolLM-135M( upped to 0.18b Use this frankenbase for training. Sorry for the mislabelling, the model is a 0.18b 181m pa...

text generationBy nisten

Qwen/Qwen2.5-14B-Instruct

apache-2.0

Qwen2.5-14B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base langu...

text generationBy Qwen

meta-llama/LlamaGuard-7b

llama2

Visit HuggingFace for more details.

text generationBy meta-llama

tiiuae/falcon-mamba-7b

other

Table of Contents 0. TL;DR(TL;DR) 1. Model Details(model-details) 2. Usage(usage) 3. Training Details(training-details) 4. Evaluation(evalua...

text generationBy tiiuae
Showing 10 of 563