SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Image classification Datasets

6 datasets in the SAVRN Model Hub for image classification, from publishers including University of Toronto Computer Science, Marco Alexandre de Oliveira Ferreira Alho, Yan, Large Scale Visual Recognition Challenge.

6 datasets.

Dataset · Image classification

typed_digital_signatures

Ben

Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size

Publicly accessible mit 10K<n<100K

Dataset · Image classification

cifar10

University of Toronto Computer Science

The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain more images from one class than another. Between them, the training batches contain exactly 5000 images from each class. - image-classification: The goal of this task is to classify a given image into one of 10 classes. The leaderboard is available here. English A…

Publicly accessible unknown 10K<n<100K

Dataset · Image classification

imagenet-1k

Large Scale Visual Recognition Challenge

ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. This dataset provides access to ImageNet (ILSVRC) 2012 which is the most commonly used subset of ImageNet. This dataset spans 1000 object classes and contains 1,281,167 training images, 50,000 validation images and 100,000 test…

Access requested at publisher other 1M<n<10M

Dataset · Image classification

Dataset

Yan

A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal pairs across 7 distinct imaging modalities, with comprehensive clinical metadata, expert lesion segmentations, and hierarchical diagnostic taxonomies. All archives are compressed into high-speed Zstandard chunks (.tar.zst) with pre-indexed SHA-256 manifests. A centralized, fully de-identified master index is provided in…

Publicly accessible cc-by-nc-sa-4.0 1M<n<10M

Dataset · Image classification

ASIMOW

AI4Manufacturing

Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first arc-welding dataset. Each record is one labelled stretch of a pulsed gas-metal arc weld, drawn as the heat input of every current pulse along that stretch with the heat-input window laid over it: a thin grey trace through the individual pulses, a rolling-median line, and a shaded band with dashed limits. The only thing to read off it is how much of the trace leaves the band. That is not the usual Category-C picture, and it is not an oversight. Every other signal dataset here asks which fault line is present; welding has none. A welder judges a stretch of weld by how much heat went into it and whether…

Access requested at publisher cc-by-4.0

Dataset · Image classification

DrsuperIA

Marco Alexandre de Oliveira Ferreira Alho

This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive or sensitive nature. There is work in progress to filter a similar dataset from winter21, and there is already ImageNet-21k-P but with different thresholds &…

Publicly accessible other 10M<n<100M

Who Publishes These Datasets

Other tasks

See all