pyannote/speaker-diarization
mitVisit HuggingFace for more details.
facebook/seamless-m4t-v2-large
cc-by-nc-4.0SeamlessM4T v2 SeamlessM4T is our foundational all-in-one Massively Multilingual and Multimodal Machine Translation model delivering high-qu...
openai/whisper-large
apache-2.0Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data...
distil-whisper/distil-large-v3
mitDistil-Whisper: distil-large-v3 Distil-Whisper was proposed in the paper Robust Knowledge Distillation via Large-Scale Pseudo Labelling( Thi...
jonatasgrosman/wav2vec2-large-xlsr-53-english
apache-2.0Fine-tuned XLSR-53 large model for speech recognition in English Fine-tuned facebook/wav2vec2-large-xlsr-53( on English using the train and ...
nvidia/canary-1b
cc-by-nc-4.0Canary 1B img { display: inline; } !Model architecture( | !Model size( | !Language( NVIDIA NeMo Canary( is a family of multi-lingual multi-t...
Systran/faster-whisper-large-v3
mitWhisper large-v3 model for CTranslate2 This repository contains the conversion of openai/whisper-large-v3( to the CTranslate2( model format....
openai/whisper-tiny
apache-2.0Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data...
openai/whisper-medium
apache-2.0Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data...
pyannote/voice-activity-detection
mitVisit HuggingFace for more details.