Salesforce/blip-image-captioning-large
bsd-3-clauseBLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation Model card for image captioning pre...
Salesforce/blip-image-captioning-base
bsd-3-clauseBLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation Model card for image captioning pre...
microsoft/trocr-base-handwritten
mitTrOCR (base-sized model, fine-tuned on IAM) TrOCR model fine-tuned on the IAM dataset( It was introduced in the paper TrOCR: Transformer-bas...
jinhybr/OCR-Donut-CORD
mitDonut (base-sized model, fine-tuned on CORD) Donut model fine-tuned on CORD. It was introduced in the paper OCR-free Document Understanding ...
microsoft/trocr-base-printed
TrOCR (base-sized model, fine-tuned on SROIE) TrOCR model fine-tuned on the SROIE dataset( It was introduced in the paper TrOCR: Transformer...
microsoft/trocr-large-printed
TrOCR (large-sized model, fine-tuned on SROIE) TrOCR model fine-tuned on the SROIE dataset( It was introduced in the paper TrOCR: Transforme...
kha-white/manga-ocr-base
apache-2.0Manga OCR Optical character recognition for Japanese text, with the main focus being Japanese manga. It uses Vision Encoder Decoder( framewo...
Salesforce/blip-image-captioning-large
bsd-3-clauseBLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation Model card for image captioning pre...
nlpconnect/vit-gpt2-image-captioning
apache-2.0nlpconnect/vit-gpt2-image-captioning This is an image captioning model trained by @ydshieh in flax ( this is pytorch version of this( The Il...
Salesforce/blip-image-captioning-base
bsd-3-clauseBLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation Model card for image captioning pre...