nvidia/canary-1b-flash
cc-by-4.0Canary 1B Flash img { display: inline; } > 🎉 NEW: Canary 1B V2 is now available! > 🌍 25 European Languages | ⏱️ Much Improved Timestamp Pr...
facebook/wav2vec2-large-960h-lv60-self
apache-2.0Wav2Vec2-Large-960h-Lv60 + Self-Training Facebook's Wav2Vec2( The large model pretrained and fine-tuned on 960 hours of Libri-Light and Libr...
openai/whisper-large-v2
apache-2.0Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data...
microsoft/Phi-4-multimodal-instruct
mit🎉Phi-4: mini-reasoning( | reasoning( | multimodal-instruct( | onnx( mini-instruct( | onnx( Model Summary Phi-4-multimodal-instruct is a lig...
openai/whisper-large-v3
apache-2.0Whisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Spee...
openai/whisper-large-v3-turbo
mitWhisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Spee...
pyannote/speaker-diarization-3.1
mitVisit HuggingFace for more details.
facebook/seamless-m4t-v2-large
cc-by-nc-4.0SeamlessM4T v2 SeamlessM4T is our foundational all-in-one Massively Multilingual and Multimodal Machine Translation model delivering high-qu...
openai/whisper-large-v2
apache-2.0Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data...
openai/whisper-large-v3
apache-2.0Whisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Spee...