SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Visual Document Retrieval Datasets

7 open-weight visual document retrieval datasets in the SAVRN Model Hub, with Massive Text Embedding Benchmark publishing the most.

7Datasets
1Publishers
1Licenses

Most Downloaded

DatasetPublisherLicenseMonthly downloads
Vidore3EnergyBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3HrBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3ComputerScienceBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3PharmaceuticalsBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3IndustrialBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3FinanceFrBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —
Vidore3FinanceEnBGEm3Reranking.v2 Massive Text Embedding Benchmark cc-by-4.0 —

Licenses

LicenseDatasetsCommercial use
cc-by-4.07Yes

Who Publishes Them

PublisherDatasets
Massive Text Embedding Benchmark7

All 7 Datasets

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you…

Publicly accessible cc-by-4.0

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Energy Fr, is a corpus of reports on energy supply in europe, intended for complex-document understanding tasks. Original queries were created in french, then translated to english, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this…

Publicly accessible cc-by-4.0

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This task, Finance - EN, is a corpus of reports from american banking companies, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use…

Publicly accessible cc-by-4.0

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This task, Finance - FR, is a corpus of reports from french companies in the luxury domain, intended for long-document understanding tasks. Original queries were created in french, then translated to english, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If…

Publicly accessible cc-by-4.0

Dataset · Visual document retrieval

Vidore3HrBGEm3Reranking.v2

Massive Text Embedding Benchmark

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, HR, is a corpus of reports released by the european union, intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this…

Publicly accessible cc-by-4.0

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Industrial reports, is a corpus of technical documents on military aircraft (fueling, mechanics...), intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out…

Publicly accessible cc-by-4.0

This task is used to evaluate the reranking performance for visual document retrieval. The candidates for reranking are the top-50 pages retrieved by the BAAI/bge-m3 model. This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish.This version add the OCR'ed markdown to allow for comparison across image-text, image-only and text-only models. You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. If you use this dataset…

Publicly accessible cc-by-4.0

Other Tasks

See all