bigcode/santacoder
bigcode-openrail-mSantaCoder !banner( Play with the model on the SantaCoder Space Demo( Table of Contents 1. Model Summary(model-summary) 2. Use(use) 3. Limit...
Qwen/Qwen3-8B
apache-2.0Qwen3-8B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense an...
Qwen/Qwen2.5-7B-Instruct-1M
apache-2.0Qwen2.5-7B-Instruct-1M Introduction Qwen2.5-1M is the long-context version of the Qwen2.5 series models, supporting a context length of up t...
stabilityai/stablelm-3b-4e1t
cc-by-sa-4.0StableLM-3B-4E1T Model Description StableLM-3B-4E1T is a 3 billion parameter decoder-only language model pre-trained on 1 trillion tokens of...
Gryphe/MythoMax-L2-13b
otherWith Llama 3 released, it's time for MythoMax to slowly fade away... Let's do it in style!( An improved, potentially even perfected variant ...
GSAI-ML/LLaDA-8B-Instruct
mitLLaDA-8B-Instruct We introduce LLaDA, a diffusion model with an unprecedented 8B scale, trained entirely from scratch, rivaling LLaMA3 8B in...
succinctly/text2image-prompt-generator
cc-by-2.0This is a GPT-2 model fine-tuned on the succinctly/midjourney-prompts( dataset, which contains 250k text prompts that users issued to the Mi...
togethercomputer/GPT-JT-6B-v1
apache-2.0GPT-JT Feel free to try out our Online Demo( Model Summary > With a new decentralized training algorithm, we fine-tuned GPT-J (6B) on 3.53 b...
EleutherAI/gpt-neo-1.3B
mitGPT-Neo 1.3B Model Description GPT-Neo 1.3B is a transformer model designed using EleutherAI's replication of the GPT-3 architecture. GPT-Ne...
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1
otherLlama-3.1-Nemotron-Ultra-253B-v1 Model Overview !Accuracy Plot(./accuracyplot.png) Llama-3.1-Nemotron-Ultra-253B-v1 is a large language mode...