SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: Some of the ios are empty. The reason is that when executing the code, the input/output sizes are too large and exceed our required constraints. Thus, they are not stored or used later. Note: Due to imperfect LLM-based transformations, some problem descriptions do not contain enough information to describe the code. We leave this as future work to further enhance our data and update it to a better version.

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: Some of the ios are empty. The reason is that when executing the code, the input/output sizes are too large and exceed our required constraints. Thus, they are not stored or used later. Note: Due to imperfect LLM-based transformations, some problem descriptions do not contain enough information to describe the code. We leave this as future work to further enhance our data and update it to a better version.

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: Some of the ios are empty. The reason is that when executing the code, the input/output sizes are too large and exceed our required constraints. Thus, they are not stored or used later. Note: Due to imperfect LLM-based transformations, some problem descriptions do not contain enough information to describe the code. We leave this as future work to further enhance our data and update it to a better version.

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: Some of the ios are empty. The reason is that when executing the code, the input/output sizes are too large and exceed our required constraints. Thus, they are not stored or used later. Note: Due to imperfect LLM-based transformations, some problem descriptions do not contain enough information to describe the code. We leave this as future work to further enhance our data and update it to a better version.

Publicly accessible

the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: "problemdescription"

Publicly accessible

We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team. The data format for each line in the 0368500filteredv2ds25.sced.jsonl is as follows: Some of the ios are empty. The reason is that when executing the code, the input/output sizes are too large and exceed our required constraints. Thus, they are not stored or used later. Note: Due to imperfect LLM-based transformations, some problem descriptions do not contain enough information to describe the code. We leave this as future work to further enhance our data and update it to a better version.

Publicly accessible

Dataset · Image classification

DrsuperIA

Marco Alexandre de Oliveira Ferreira Alho

This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive or sensitive nature. There is work in progress to filter a similar dataset from winter21, and there is already ImageNet-21k-P but with different thresholds &…

Publicly accessible other 10M<n<100M

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.