# Visual Document Retrieval Datasets: AI Datasets
Source: https://savrn.com/datasets/tasks/visual-document-retrieval
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

SAVRN Model Hub · Datasets by Task

# Visual Document Retrieval Datasets

7 open-weight visual document retrieval datasets in the SAVRN Model Hub, with Massive Text Embedding Benchmark publishing the most.

7Datasets

1Publishers

1Licenses

## Most Downloaded

| Dataset | Publisher | License | Monthly downloads |
| --- | --- | --- | --- |
| [Vidore3EnergyBGEm3Reranking.v2](https://savrn.com/datasets/vidore3energybgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3HrBGEm3Reranking.v2](https://savrn.com/datasets/vidore3hrbgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3ComputerScienceBGEm3Reranking.v2](https://savrn.com/datasets/vidore3computersciencebgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3PharmaceuticalsBGEm3Reranking.v2](https://savrn.com/datasets/vidore3pharmaceuticalsbgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3IndustrialBGEm3Reranking.v2](https://savrn.com/datasets/vidore3industrialbgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3FinanceFrBGEm3Reranking.v2](https://savrn.com/datasets/vidore3financefrbgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |
| [Vidore3FinanceEnBGEm3Reranking.v2](https://savrn.com/datasets/vidore3financeenbgem3reranking-v2) | Massive Text Embedding Benchmark | cc-by-4.0 | — |

## Licenses

| License | Datasets | Commercial use |
| --- | --- | --- |
| cc-by-4.0 | 7 | Yes |

## Who Publishes Them

| Publisher | Datasets |
| --- | --- |
| [Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb) | 7 |

## All 7 Datasets

Dataset · Visual document retrieval

### [Vidore3ComputerScienceBGEm3Reranking.v2](https://savrn.com/datasets/vidore3computersciencebgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3computersciencebgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3EnergyBGEm3Reranking.v2](https://savrn.com/datasets/vidore3energybgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Energy Fr, is a corpus of reports on energy supply in europe, intended for complex-document understanding tasks. Original queries were created in french, then translated to english, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3energybgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3FinanceEnBGEm3Reranking.v2](https://savrn.com/datasets/vidore3financeenbgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This task, Finance - EN, is a corpus of reports from american banking companies, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3financeenbgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3FinanceFrBGEm3Reranking.v2](https://savrn.com/datasets/vidore3financefrbgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This task, Finance - FR, is a corpus of reports from french companies in the luxury domain, intended for long-document understanding tasks. Original queries were created in french, then translated to english, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3financefrbgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3HrBGEm3Reranking.v2](https://savrn.com/datasets/vidore3hrbgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, HR, is a corpus of reports released by the european union, intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3hrbgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3IndustrialBGEm3Reranking.v2](https://savrn.com/datasets/vidore3industrialbgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Industrial reports, is a corpus of technical documents on military aircraft (fueling, mechanics...), intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3industrialbgem3reranking-v2)

Dataset · Visual document retrieval

### [Vidore3PharmaceuticalsBGEm3Reranking.v2](https://savrn.com/datasets/vidore3pharmaceuticalsbgem3reranking-v2)

[Massive Text Embedding Benchmark](https://savrn.com/model-publishers/mteb)

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this dataset…

Publicly accessible cc-by-4.0

[View dataset](https://savrn.com/datasets/vidore3pharmaceuticalsbgem3reranking-v2)

## Other Tasks

- [Text Generation](https://savrn.com/datasets/tasks/text-generation) 96
- [Robotics](https://savrn.com/datasets/tasks/robotics) 82
- [Text Classification](https://savrn.com/datasets/tasks/text-classification) 38
- [Other](https://savrn.com/datasets/tasks/other) 34
- [Text Retrieval](https://savrn.com/datasets/tasks/text-retrieval) 33
- [Question Answering](https://savrn.com/datasets/tasks/question-answering) 31
- [Speech Recognition](https://savrn.com/datasets/tasks/speech-recognition) 23
- [Time Series Forecasting](https://savrn.com/datasets/tasks/time-series-forecasting) 18
- [Image Classification](https://savrn.com/datasets/tasks/image-classification) 15
- [Image to Text](https://savrn.com/datasets/tasks/image-to-text) 14
- [Tabular Classification](https://savrn.com/datasets/tasks/tabular-classification) 12
- [Visual Question Answering](https://savrn.com/datasets/tasks/visual-question-answering) 12

[See all](https://savrn.com/datasets/tasks)
