Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [\[checkpoint\]](https://huggingface.co/wyu1/Leopard-LLaVA) and Leopard-Idefics2 [\[checkpoint\]](https://huggingface.co/wyu1/Leopard-Idefics2). - to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0) - to load all the subsets of the images