SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Video text to text Datasets

2 datasets in the SAVRN Model Hub for video text to text, from publishers including MVP Lab, LMMs-Lab.

2 datasets.

Dataset · Video text to text

LLaVA-OneVision-2-Data

MVP Lab

Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. The dataset is split across two Hugging Face repositories because of its size: If you need the long-video data, download the video shards from Part 2 and use the corresponding captions and mapping files from Part 1. The directory names in Part 2 are upload prefixes. Their duration mapping is: The mapping above was verified from the sample paths stored inside the WebDataset archives. For the ~180-second split: 1. Download video shards from tom/ in Part 2. 2. Read captions…

Publicly accessible apache-2.0 10M<n<100M

Dataset · Video text to text

EgoLife

LMMs-Lab

Data cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info. Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. Code: https://github.com/egolife-ai/EgoLife

Publicly accessible mit

Who Publishes These Datasets

Other tasks

See all