Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 154 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7345–7392 of 12,172
Acticipate is a publicly available dataset with recordings of human body-motion and eye-gaze, acquired in an experimental scenario with an actor interacting with three subjects.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ActioNet is a video task-based dataset collected in a synthetic 3D environment.
1 paper · 0 benchmarks
This is a public dataset for evaluating multi-object detection algorithms in active Terahertz imaging resolution 5 mm by 5 mm.
1 paper · 0 benchmarks
Consists of 10,000+ video-sentence pairs with each accompanied by an annotated sentence specified video thumbnail.
1 paper · 0 benchmarks
AdobeIndoorNav is a dataset collected in real-world to facilitate the research in DRL based visual navigation.
1 paper · 0 benchmarks
AdvSuffixes - Information AdvSuffixes is a curated dataset of adversarial prompts and suffixes designed to evaluate and enhance the robustness of large language models (LLMs) against adversarial attacks.
1 paper · 0 benchmarks
The Advice-Seeking Questions (ASQ) dataset is a collection of personal narratives with advice-seeking questions.
1 paper · 0 benchmarks
We filter and match the landmarks in the Google Landmarks dataset with their OpenStreetMap polygons and filter for those located in the United States, resulting in 602 landmarks.
1 paper · 0 benchmarks
AerialMPT is a dataset for pedestrian tracking in aerial image sequences and presents real-world challenges for MOT algorithms such as low frame rate, small moving objects, and complex backgrounds.
1 paper · 0 benchmarks
AesVQA is a dataset that contains 72168 high-quality images and 324756 pairs of aesthetic questions.
1 paper · 0 benchmarks
AfriSenti (AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages)
AfriSenti is the largest sentiment analysis dataset for under-represented African languages, covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican…
1 paper · 1 benchmark
EEG signals from 60 users have been recorded whose age range lies between 6 and 55 years.
1 paper · 0 benchmarks
AgentEval is part of the AgentGym framework, which is designed to evaluate and develop generally-capable Large Language Model-based (LLM-based) agents.
1 paper · 0 benchmarks
AgroLens (AgroLens Soil Prediction Dataset)
This dataset has been curated for a student research project at the Technische Hochschule Ingolstadt with Mi4Poeople and its Soil project (https://de.mi4people.org/soil-quality-evaluation-system).
1 paper · 0 benchmarks
This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly.
1 paper · 0 benchmarks
Synthetic Dataset created in AirSim
1 paper · 0 benchmarks
AjwaOrMedjool (AjwaOrMedjool: a binary balanced dataset to teach machine learning)
The dataset contains three subsets: 1- a dataset containing hand-crafted features to classify two types of organic dates (Ajwa or Medjool); 2- a dataset containing tabular data with features created automatically using deep learning to…
1 paper · 0 benchmarks
Sentiment Analysis of Movie Reviews in Albanian
1 paper · 0 benchmarks
Named Entity Recognition in Albanian
1 paper · 0 benchmarks
a corpus for topic modeling in Albanian
1 paper · 0 benchmarks
Alex-20: contains ~1.3M general inorganic materials curated from the Alexandria database, with energy above the convex hull less than 0.1 eV/atom and no more than 20 atoms in unit cell.
1 paper · 0 benchmarks
This dataset is composed of the URLs of the top 1 million websites.
1 paper · 0 benchmarks
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
We introduce the novel task of multimodal puzzle solving, framed within the context of visual question-answering.
1 paper · 1 benchmark
The Algonauts 2023 Challenge focuses on predicting responses in the human brain as participants perceive complex natural visual scenes.
1 paper · 0 benchmarks
All Conference Alert is a tech startup dedicated to helping organizers and attendees of academic conferences, seminars, workshops, and webinars.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The AllMusic Mood Subset (AMS) is a dataset for mood classification from songs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset contains sentences from Amazon customer reviews (sampled from Amazon product review dataset) annotated for counterfactual detection (CFD) binary classification.
1 paper · 0 benchmarks
The Amazon Polarity dataset is a set of reviews from Amazon.
1 paper · 1 benchmark
The Ambiguous VQA dataset is a dataset of ambiguous questions about images.
1 paper · 0 benchmarks
Amharic Error Corpus is a manually annotated spelling error corpus for Amharic, lingua franca in Ethiopia.
1 paper · 0 benchmarks
Among Them (Among Them dialogs and persuasion labels)
The dataset contains dialogs of different LLMs from the discussion phase of a text-based Among Us-like game.
1 paper · 0 benchmarks
This R package, documented in a very similar way to the book R4DS, provides functions to replicate the original Stata results from the book An Advanced Guide to Trade Policy Analysis.
1 paper · 0 benchmarks
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable.
1 paper · 1 benchmark
Data collection: Finding a suitable source of data is considered a first step toward building a database.
1 paper · 1 benchmark
This repository contains the datasets corresponding to the three benchmark problems for the Fourier Neural Mappings scientific machine learning architectures.
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
Analytic provenance is a data repository that can be used to study human analysis activity, thought processes, and software interaction with visual analysis tools during exploratory data analysis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset was constructed from an analysis of about 1.5 million apps from Google Play to identify a set of common libraries, to facilitate Android app analysis.
1 paper · 0 benchmarks
The AneuX morphology database includes data from 3 different data sources: AneuX, @neurIST and Aneurisk.
1 paper · 0 benchmarks
The Angry Tweets dataset is a collection of anonymized Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.