Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 20 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 913–960 of 12,172
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
The CAD-60 and CAD-120 data sets comprise of RGB-D video sequences of humans performing activities which are recording using the Microsoft Kinect sensor.
65 papers · 1 benchmark
CALVIN (Composing Actions from Language and Vision)
CALVIN (Composing Actions from Language and Vision), is an open-source simulated benchmark to learn long-horizon language-conditioned robot manipulation tasks.
65 papers · 2 benchmarks
DBP15k contains four language-specific KGs that are respectively extracted from English (En), Chinese (Zh), French (Fr) and Japanese (Ja) DBpedia, each of which contains around 65k-106k entities.
65 papers · 3 benchmarks
DuReader is a large-scale open-domain Chinese machine reading comprehension dataset.
65 papers · 1 benchmark
FakeAVCeleb is a novel Audio-Video Deepfake dataset that not only contains deepfake videos but respective synthesized cloned audios as well.
65 papers · 2 benchmarks
Friendster is an on-line gaming network.
65 papers · 0 benchmarks
MMI (MMI Facial Expression Database)
The MMI Facial Expression Database consists of over 2900 videos and high-resolution still images of 75 subjects.
65 papers · 1 benchmark
MagnaTagATune dataset contains 25,863 music clips.
65 papers · 2 benchmarks
The Newer College Dataset is a large dataset with a variety of mobile mapping sensors collected using a handheld device carried at typical walking speeds for nearly 2.2 km through New College, Oxford.
65 papers · 0 benchmarks
Occluded REID is an occluded person dataset captured by mobile cameras, consisting of 2,000 images of 200 occluded persons (see Fig.
65 papers · 1 benchmark
PASCAL-Part is a set of additional annotations for PASCAL VOC 2010.
65 papers · 4 benchmarks
The Places365 dataset is a scene recognition dataset.
65 papers · 7 benchmarks
This dataset consists of (human-written) NBA basketball game summaries aligned with their corresponding box- and line-scores.
65 papers · 5 benchmarks
Wildtrack is a large-scale and high-resolution dataset.
65 papers · 2 benchmarks
AIDA CoNLL-YAGO contains assignments of entities to the mentions of named entities annotated for the original CoNLL 2003 entity recognition task.
64 papers · 0 benchmarks
AQUA-RAT (Algebra Question Answering with Rationales)
Algebra Question Answering with Rationales (AQUA-RAT) is a dataset that contains algebraic word problems with rationales.
64 papers · 0 benchmarks
ASTE (Aspect Sentiment Triplet Extraction)
Target-based sentiment analysis or aspect-based sentiment analysis (ABSA) refers to addressing various sentiment analysis tasks at a fine-grained level, which includes but is not limited to aspect extraction, aspect sentiment…
64 papers · 1 benchmark
The EmpatheticDialogues dataset is a large-scale multi-turn empathetic dialogue dataset collected on the Amazon Mechanical Turk, containing 24,850 one-to-one open-domain conversations.
64 papers · 2 benchmarks
HHH (Helpful, Honest, & Harmless)
The HHH dataset, also known as the Helpful, Honest, & Harmless (HHH) Alignment dataset, is a dataset used for evaluating language models.
64 papers · 0 benchmarks
MTOP (Multilingual Task-Oriented Semantic Parsing)
A multilingual task-oriented semantic parsing dataset covering 6 languages and 11 domains.
64 papers · 0 benchmarks
MagicBrush is a manually-annotated instruction-guided image editing dataset covering diverse scenarios single-turn, multi-turn, mask-provided, and mask-free editing.
64 papers · 0 benchmarks
The Mall is a dataset for crowd counting and profiling research.
64 papers · 1 benchmark
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
AgentBench is a comprehensive benchmark designed to evaluate Large Language Models (LLMs) as agents in interactive environments.
63 papers · 0 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
ESD (Emotional Speech Database)
ESD is an Emotional Speech Database for voice conversion research.
63 papers · 0 benchmarks
LFWA is a popular unconstrained facial attribute dataset, which consists of 13,143 facial images of 5,749 identities.
63 papers · 1 benchmark
LRS3-TED is a multi-modal dataset for visual and audio-visual speech recognition.
63 papers · 7 benchmarks
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
SBIC (Social Bias Inference Corpus)
To support large-scale modelling and evaluation with 150k structured annotations of social media posts, covering over 34k implications about a thousand demographic groups.
63 papers · 0 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
Contains 51,583 descriptions of 11,046 objects from 800 ScanNet scenes.
63 papers · 1 benchmark
SceneNN is an RGB-D scene dataset consisting of more than 100 indoor scenes.
63 papers · 1 benchmark
Sentence Compression is a dataset where the syntactic trees of the compressions are subtrees of their uncompressed counterparts, and hence where supervised systems which require a structural alignment between the input and output can be…
63 papers · 0 benchmarks
Atari Games for only 100k environment steps.
62 papers · 0 benchmarks
COCO-QA is a dataset for visual question answering.
62 papers · 0 benchmarks
ComplexWebQuestions is a dataset for answering complex questions that require reasoning over multiple web snippets.
62 papers · 2 benchmarks
Dataset containing Credit scores and loan repayment rate (90-day default rate) for individuals, separated by race (white, black, Hispanic Asian).
62 papers · 0 benchmarks
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
FaceScape dataset provides 3D face models, parametric models and multi-view images in large-scale and high-quality.
62 papers · 1 benchmark
Gaze360 (Physically Unconstrained Gaze Estimation in the Wild)
Understanding where people are looking is an informative social cue.
62 papers · 1 benchmark
This dataset contains 21,889 outfits from polyvore.com, in which 17,316 are for training, 1,497 for validation and 3,076 for testing.
62 papers · 3 benchmarks
TNL2K (Tracking by natural language)
Tracking by Natural Language (TNL2K) is constructed for the evaluation of tracking by natural language specification.
62 papers · 2 benchmarks
TaxiBJ consists of trajectory data from taxicab GPS data and meteorology data in Beijing from four time intervals: 1st Jul.
62 papers · 3 benchmarks
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
62 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.