Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 166 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7921–7968 of 12,172
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
The 50-ha plot at Barro Colorado Island was initially demarcated and fully censused in 1982, and has been fully censused 7 times since, every 5 years from 1985 through 2015.
1 paper · 0 benchmarks
In everyday language processing, sentence context affects how readers and listeners process upcoming words.
1 paper · 0 benchmarks
Experimental setup (learner code, data generator) for comparing different sequence processors on a dataset generated from motion variables of a pendulum with exponentially-increasing string length.
1 paper · 0 benchmarks
The Complex-TV-QA dataset, to our knowledge, is the inaugural resource that provides human-annotated, detailed video captions within traffic scenarios, alongside complex reasoning questions.
1 paper · 0 benchmarks
The Composed Quora dataset consists of questions extracted from Quora that are grouped together if they are asking the same thing.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We capture some hyperspectral images in our lab using the multishot DD-CASSI architecture.
1 paper · 0 benchmarks
The computer codes (in Matlab or Fortran) can be downloaded from the website: people.maths.ox.ac.uk/erban/Education/
1 paper · 0 benchmarks
Computer Vision Arxiv Figures dataset consists of 88,645 images that more closely resemble the structure of our visual prompts.
1 paper · 0 benchmarks
This is a corpus of about 500 computer vision datasets, from which the authors sampled 114 dataset publications across different vision tasks and coded for themes through both structured and qualitative content analysis.
1 paper · 0 benchmarks
ConCon: Continually Confounded Dataset is a confounded visual dataset for continual learning.
1 paper · 0 benchmarks
ConQA (Conceptual Query Answering)
ConQA is a dataset created using the intersection between VisualGenome and MS-COCO.
1 paper · 2 benchmarks
ConSLAM (Construction Dataset for SLAM)
ConSLAM is a real-world dataset collected periodically on a construction site to measure the accuracy of mobile scanners' SLAM algorithms.
1 paper · 0 benchmarks
The ConScenD dataset consists of over 340 scenarios extracted from the naturalistic highway dataset highD.
1 paper · 0 benchmarks
Concept-1K contains 1023 novel concepts from six domains, including economy, culture, science and technology, environment, education, and health and medical.
1 paper · 0 benchmarks
Concrete is the most important material in civil engineering.
1 paper · 1 benchmark
We construct a large-scale conducting motion dataset, named ConductorMotion100, by deploying pose estimation on conductor view videos of concert performance recordings collected from online video platforms.
1 paper · 0 benchmarks
This is a video and image segmentation dataset for human head and shoulders, relevant for creating elegant media for videoconferencing and virtual reality applications.
1 paper · 0 benchmarks
French sentences are sourced from Tatoeba repository and then translated into Congolese Swahili.
1 paper · 0 benchmarks
Conic10K is an open-ended math problem dataset on conic sections in Chinese senior high school education.
1 paper · 0 benchmarks
ConsInv is a stereo RGB + IMU dataset designed for Dynamic SLAM testing and contains two subsets: - ConsInv-Indoors contains sequences in an office setting where small objects are moved.
1 paper · 0 benchmarks
Description - Repository: Code, Page, Data - Paper: arxiv.org/abs/2411.17440 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a star and citation.
1 paper · 0 benchmarks
Enumerate–Conjecture–Prove: Formally Solving Answer-Construction Problem in Math Competitions We release the ConstructiveBench dataset as part of our Enumerate–Conjecture–Prove (ECP) paper.
1 paper · 0 benchmarks
The progress of Large Language Models (LLMs) has largely been driven by the availability of large-scale unlabeled text data for unsupervised learning.
1 paper · 0 benchmarks
This is a revised and extended second version of a Contextualised Polyseme Word Sense Dataset.
1 paper · 0 benchmarks
Continuous Defect Prediction (CDP) is a dataset of more than 11 million data rows, representing files involved in Continuous Integration (CI) builds, that synthesize the results of CI builds with data mined from software repositories.
1 paper · 0 benchmarks
Corpus of controversial news articles extracted from Twitter.
1 paper · 0 benchmarks
ConvSumX is a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context.
1 paper · 0 benchmarks
This dataset contains 12,500 meter images acquired in the field by the employees of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units, across 395 cities and 1,113 locations (i.e., districts,…
1 paper · 1 benchmark
CoreSearch is a dataset for Cross-Document Event Coreference Search.
1 paper · 0 benchmarks
CoronaVis is a dataset of tweets related to coronavirus.
1 paper · 0 benchmarks
This is a dataset for coronavirus-themed malware for Android devices.
1 paper · 0 benchmarks
The dataset has thousands of time series.
1 paper · 0 benchmarks
Correlated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation.
1 paper · 0 benchmarks
This original data set includes the following four sheets: Sheet 1: Raw Data (the original data set) Sheet 2: Variables (A list with the variables included in the study) Sheet 3: Countries Scientific Relative Production Sheet 4:…
1 paper · 0 benchmarks
Dataset of Stack Overflow questions about Terraform with cost-related keywords.
1 paper · 0 benchmarks
Dataset of commit messages and issues containing evidence of cost awareness.
1 paper · 0 benchmarks
Using Council Data Project infrastructures (https://councildataproject.org), we assemble longitudinal municipal council meeting transcript data.
1 paper · 0 benchmarks
Probing cross-modal capabilities of Vision & Language models with a counting task.
1 paper · 0 benchmarks
Higher education plays a critical role in driving an innovative economy by equipping students with knowledge and skills demanded by the workforce.
1 paper · 0 benchmarks
The Courtois NeuroMod project aims at training artificial neural networks using extensive experimental data on individual human brain activity and behaviour.
1 paper · 0 benchmarks
CoverageEval is a dataset specifically designed for evaluating LLMs on this task.
1 paper · 0 benchmarks
CovidET-Appraisals is the most comprehensive dataset to-date that assesses 24 cognitive appraisal dimensions of emotions, each with a natural language rationale, across 241 Reddit posts.
1 paper · 0 benchmarks
CovidET-EXT is a dataset that augments Zhan et al.
1 paper · 0 benchmarks
This Dataset contains pairs off textual natural language questions and SPARQL queries on a small subset of the CoyPu KnowledgeGraph (https://coypu.org/ergebnisse/knowledge-graph)
1 paper · 0 benchmarks
Survey responses where all creative habit ordinal responses were converted to Creative Habit Tags - these tags were used in the analysis to build a network of people linked if they share similar creative habit sets, or a network of…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.