Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. The dataset is split across two Hugging Face repositories because of its size: If you need the long-video data, download the video shards from Part 2 and use the corresponding captions and mapping files from Part 1. The directory names in Part 2 are upload prefixes. Their duration mapping is: The mapping above was verified from the sample paths stored inside the WebDataset archives. For the ~180-second split: 1. Download video shards from tom/ in Part 2. 2. Read captions…
