Dataset · Image classification
Ben
Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size
Publicly accessible
mit
10K<n<100K
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain more images from one class than another. Between them, the training batches contain exactly 5000 images from each class. - image-classification: The goal of this task is to classify a given image into one of 10 classes. The leaderboard is available here. English A…
Publicly accessible
unknown
10K<n<100K
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. This dataset provides access to ImageNet (ILSVRC) 2012 which is the most commonly used subset of ImageNet. This dataset spans 1000 object classes and contains 1,281,167 training images, 50,000 validation images and 100,000 test…
Access requested at publisher
other
1M<n<10M
Y
Dataset · Image classification
Yan
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal pairs across 7 distinct imaging modalities, with comprehensive clinical metadata, expert lesion segmentations, and hierarchical diagnostic taxonomies. All archives are compressed into high-speed Zstandard chunks (.tar.zst) with pre-indexed SHA-256 manifests. A centralized, fully de-identified master index is provided in…
Publicly accessible
cc-by-nc-sa-4.0
1M<n<10M
Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first arc-welding dataset. Each record is one labelled stretch of a pulsed gas-metal arc weld, drawn as the heat input of every current pulse along that stretch with the heat-input window laid over it: a thin grey trace through the individual pulses, a rolling-median line, and a shaded band with dashed limits. The only thing to read off it is how much of the trace leaves the band. That is not the usual Category-C picture, and it is not an oversight. Every other signal dataset here asks which fault line is present; welding has none. A welder judges a stretch of weld by how much heat went into it and whether…
Access requested at publisher
cc-by-4.0
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive or sensitive nature. There is work in progress to filter a similar dataset from winter21, and there is already ImageNet-21k-P but with different thresholds &…
Publicly accessible
other
10M<n<100M