Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 241 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11521–11568 of 12,172
Eduge (Eduge news classification dataset)
Eduge news classification dataset provided by Bolorsoft LLC.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 5000+ original Electronic Items images captured and crowdsourced from over 1000+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
EmoNoBa (EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts)
Detecting Multi-labeled Emotion for 6 emotion categories, namely Love, Joy, Surprise, Anger, Sadness, Fear.
0 papers · 0 benchmarks
The English-Pashto Language Dataset (EPLD) is a comprehensive resource aimed to provide linguistic insights into the Pashto language.
0 papers · 0 benchmarks
Briefly describe the dataset.
0 papers · 0 benchmarks
This is a who-trust-whom online social network of a general consumer review site Epinions.com.
0 papers · 0 benchmarks
A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks.
0 papers · 0 benchmarks
Extended BBC Pose is a pose estimation dataset which extends the BBC Pose dataset with 72 additional training videos.
0 papers · 0 benchmarks
This COVID-19 dataset consists of Non-COVID and COVID cases of both X-ray and CT images.
0 papers · 0 benchmarks
This is a machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS train set.
0 papers · 0 benchmarks
This is an improved machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS [1] set.
0 papers · 0 benchmarks
Description The F1 Regulations, Safety, and Racing Performance dataset provides an overview of key factors influenced by the evolving Fédération Internationale de l'Automobile (FIA) regulations from 1990 to 2023.
0 papers · 0 benchmarks
FALLMUD (FAscicle Lower Leg Muscle Ultrasound Dataset)
FAscicle Lower Leg Muscle Ultrasound Dataset is a dataset composed of 812 ultrasound images of lower leg muscles to analyze muscle weaknesses and prevent injuries.
0 papers · 0 benchmarks
FHRMA is an open-source project for Fetal Heart Rate Morphological Analysis containing Matlab source code and datasets.
0 papers · 0 benchmarks
This is the list of datasets used for Conti's FNS-Funded projects
0 papers · 0 benchmarks
FPV-O is a multi-subject first-person vision dataset of office activities.
0 papers · 0 benchmarks
FRIDA (Foggy Road Image Database)
FRIDA and FRIDA2 are databases of numerical images easily usable to evaluate in a systematic way the performance of visibility and contrast restoration algorithms.
0 papers · 0 benchmarks
FSI (Fluid-Solid interaction)
Data Set Structure Fluid Structure Interaction(NS +Elastic wave) The TFfsi2results folder contains simulation data organized by various parameters (mu, x1, x2) where mu determines the viscosity and x1 and x2 are the parameters of the inlet…
0 papers · 0 benchmarks
FSL4 (Freesound Loops 4k)
The FSL4 dataset contains ~4000 user-contributed loops uploaded to Freesound.
0 papers · 0 benchmarks
The FaQuAD dataset is a reading comprehension dataset designed for evaluating question-answering models.
0 papers · 0 benchmarks
The Fabrics Dataset consists of about 2000 samples of garments and fabrics.
0 papers · 0 benchmarks
The free Face dataset made for students and teachers.
0 papers · 0 benchmarks
We introduce an annotated dataset of five thousand human labeled pareidolic face images, called Faces in Things''.
0 papers · 0 benchmarks
The Family101 dataset is the a large-scale dataset of families across several generations.
0 papers · 0 benchmarks
The Fields2Benhmark dataset is a collection of 350 agricultural fields in vector format manually selected to test agricultural coverage path planning algorithms.
0 papers · 0 benchmarks
FinArg (Financial Argument Mining)
With the goal of reasoning on the financial textual data, we present a novel dataset for annotating arguments, their components, and relations in the transcripts of earnings conference calls (ECCs).
0 papers · 0 benchmarks
The FlareReal600 is a nighttime flare removal dataset, which contains 650 real-captured images pairs and 500 flare images.
0 papers · 0 benchmarks
A high-quality dataset forms the foundation for machine learning-based predictions of structural load capacity.
0 papers · 0 benchmarks
- Marine wastes are severely threatening marine animals and their habitat, also causing an impact on human life through toxic substances transportation and accumulation.
0 papers · 0 benchmarks
This dataset is being constructed specifically to support research on techniques that bridge the gap between 2D, appearance-based recognition techniques, and fully 3D approaches.
0 papers · 0 benchmarks
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
0 papers · 0 benchmarks
FluencyBank is a shared database for the study of fluency development.
0 papers · 0 benchmarks
Simulated GFP-actin-stained A549 Lung Cancer cells embedded in a Matrigel matrix Dr.
0 papers · 0 benchmarks
Developing Tribolium Castaneum embryo (3D cartographic projection) Dr.
0 papers · 0 benchmarks
The Follicular-Segmentation dataset consists of 6900 cropped typical image patches of 1024x1024 pixels containing: follicular areas, colloid areas, and the other blank background areas.
0 papers · 0 benchmarks
This dataset was created for Fongbe automatic speech recognition task and contains about 3979 recordings of 13 participants reading a text written in Fongbe, one sentence at a time.
0 papers · 0 benchmarks
ForeDeCk is a time series database compiled at the National Technical University of Athens that contains 900,000 continuous time series, built from multiple, diverse and publicly accessible sources.
0 papers · 0 benchmarks
We introduce FortisAVQA, a dataset designed to assess the robustness of AVQA models.
0 papers · 0 benchmarks
This dataset contains 16,000 images of four shapes; square, star, circle, and triangle.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.