Donut model fine-tuned on DocVQA. It was introduced in the paper OCR-free Document Understanding Transformer by Geewok et al. and first released in this repository. Disclaimer: The team releasing Donut did not write a model card for this model so this model card has been written by the Hugging Face team. Donut consists of a vision encoder (Swin Transformer) and a text decoder (BART). Given an image, the encoder first encodes the image into a tensor of embeddings (of shape batchsize, seqlen, hiddensize), after which the decoder autoregressively generates text, conditioned on the encoding of the encoder. This model is fine-tuned on DocVQA, a document visual question answering dataset. We…
SAVRN Model Hub · Models by Task
Document Question Answering Models
1 open-weight document question answering models in the SAVRN Model Hub, with NAVER CLOVA INFORMATION EXTRACTION publishing the most.
1Models
1Publishers
1Licenses
Most Downloaded
| Model | Publisher | Parameters | License | Monthly downloads | Cheapest GPUs at 16-bit |
|---|---|---|---|---|---|
| donut-base-finetuned-docvqa | NAVER CLOVA INFORMATION EXTRACTION | — | mit | 24.3k | — |
Licenses
| License | Models | Commercial use |
|---|---|---|
| mit | 1 | Yes |
Who Publishes Them
| Publisher | Models |
|---|---|
| NAVER CLOVA INFORMATION EXTRACTION | 1 |
All 1 Models
Questions
Which Document question answering models are most downloaded?
By monthly downloads reported by the Hugging Face Hub: donut-base-finetuned-docvqa (24.3k).