Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 178 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8497–8544 of 12,172
Composed of judgments from the French Court of cassation and their corresponding summaries.
1 paper · 0 benchmarks
We collated subcorpora each between 50,000 and 70,000 words, containing samples of national dialects of French across different countries: Algeria, Democratic Republic of Congo, France, Ivory Coast, Morocco and Senegal.
1 paper · 0 benchmarks
This dataset contains the publication data underlying the French Open Science Monitor.
1 paper · 0 benchmarks
The FrodoBots 2K Dataset is a diverse collection of camera footage, GPS, IMU, audio recordings & human control data collected from ~2,000 hours of tele-operated sidewalk robots driving in 10+ cities.
1 paper · 0 benchmarks
Fruits Dataset for Classification About Dataset (strawberries, peaches, pomegranates) Photo requirements: 1-White background 2-.jpg 3- Image size 300300 The number of photos required is 250 photos of each fruit when it is fresh and 250…
1 paper · 0 benchmarks
- The dataset contains full-spectral autofluorescence lifetime microscopic images (FS-FLIM) acquired on unstained ex-vivo human lung tissue, where 100 4D hypercubes of 256x256 (spatial resolution) x 32 (time bins) x 512 (spectral channels…
1 paper · 0 benchmarks
FullTextPeerRead is a dataset created by Jeong et al.
1 paper · 0 benchmarks
FunKPoint is a dataset for finding correspondences in visual data that has ground truth correspondences for 10 tasks and 20 object categories.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Demonstration data for 4 FurnitureBench tasks collected with a SpaceMouse using a DiffIK Controller.
1 paper · 0 benchmarks
This dataset collection provides trajectory datasets of - synthetic trajectories with variants of different noise and novelty levels - brightkite with labels for evaluation - amazon drivers with labels for evaluation - Deutsche Bahn…
1 paper · 0 benchmarks
This dataset was created to test whether it's possible to build a general-purpose detector that can tell real images apart from fake ones generated by convolutional neural networks (CNNs), no matter which model or dataset was used to…
1 paper · 0 benchmarks
GASP is a dataset composed by a list of cited abstracts associated with the corresponding source abstract.
1 paper · 0 benchmarks
GCN inference on NeuraChip accelerator on Cora dataset.
1 paper · 0 benchmarks
GD-NLI (Generated Debiased NLI Datasets)
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets.
1 paper · 0 benchmarks
The GDIT Aerial Airport dataset consists of aerial images containing instances of parked airplanes.
1 paper · 0 benchmarks
GE852 is a dataset of 852 game engine repositories mined from GitHub in two languages, namely Java and C++.
1 paper · 0 benchmarks
GED (Gridded Establishment Dataset)
GED is a dataset on the economic activity of mainland China, which measures the volume of establishments at a 0.01 latitude by 0.01 longitude scale.
1 paper · 0 benchmarks
GEM (A General Evaluation Benchmark on Multi-modal Tasks) is a significant benchmark dataset designed to evaluate the performance of cross-modal pre-trained models, including both understanding and generation tasks.
1 paper · 0 benchmarks
GEM HOUSE OPENDATA (GERMAN ELECTRICITY CONSUMPTION IN MANY HOUSEHOLDS OVER THREE YEARS 2018-2020 (FRESH ENERGY))
Filip Milojkovic, August 13, 2021, "GEM HOUSE openData: German Electricity consumption in Many HOUSEholds over three years 2018-2020 (Fresh Energy)", IEEE Dataport, doi: https://dx.doi.org/10.21227/4821-vf03.
1 paper · 0 benchmarks
Dataset Card for Dataset Name This dataset is a filtered version of BookCorpus containing only gender-neutral words.
1 paper · 0 benchmarks
GENTER (GEnder Name TEmplates with pRonouns)
This dataset consists of template sentences associating first names ([NAME]) with third-person singular pronouns ([PRONOUN]), e.g., [NAME] asked , not sounding as if [PRONOUN] cared about the answer .
1 paper · 0 benchmarks
This dataset contains short sentences linking a first name, represented by the template mask [NAME], to stereotypical associations.
1 paper · 0 benchmarks
This repository is an extension of GEval.
1 paper · 0 benchmarks
GF-PA66 3D XCT (Glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT)))
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
1 paper · 1 benchmark
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
1 paper · 0 benchmarks
GIE-Bench is a benchmark designed to evaluate text-guided image editing models across two critical dimensions: Functional correctness — assessed via VQA-style multiple-choice questions Content preservation — evaluated through object-aware…
1 paper · 0 benchmarks
The released GIF Reply dataset contains 1,562,701 real text-GIF conversation turns on Twitter.
1 paper · 1 benchmark
A random sample of 200 machine learning publications, systematically analyzed by a team of labelers, who asked up to 15 questions about how the publication discusses its training data.
1 paper · 0 benchmarks
GJ (gastrojejunostomy utsw)
49 videos of gastrojejunostomy procedure
1 paper · 0 benchmarks
GLAMI-1M (A Multilingual Image-Text Fashion Dataset)
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark.
1 paper · 1 benchmark
GLARE (Guided LexRank for Advanced Retrieval in Legal Analysis)
The Guided Lexrank algorithm is applied to dataset specialappeal.csv to summarize the texts of legal documents.
1 paper · 0 benchmarks
GLARE is an Arabic Apps Reviews dataset collected from Saudi Google PlayStore.
1 paper · 0 benchmarks
data/images: data/images/Base : 132 screenshots of game1 & game2 with UI display issues from 466 test reports.
1 paper · 0 benchmarks
A dataset for medical consultation dialogues.
1 paper · 0 benchmarks
The GMDCSA dataset contains 16 ADL (not fall) activities and 16 Fall activities.
1 paper · 0 benchmarks
GMSC (Give Me Some Credit)
Data for a Kaggle competition Banks play a crucial role in market economies.
1 paper · 0 benchmarks
GO21 is a biomedical knowledge graph that models genes, proteins, drugs, and the hierarchy of the biological processes they participate in.
1 paper · 1 benchmark
GOD (Generic Object Decoding)
The Generic Object Decoding (GOD) Dataset is a specialized resource developed for fMRI-based decoding.
1 paper · 1 benchmark
brain-image-text trimodal datasets
1 paper · 0 benchmarks
GOTOV (Human Physical Activity and Energy Expenditure Dataset on Older Individuals)
Stylianos ParaschiakosStylianos Paraschiakos, Beekman M.
1 paper · 0 benchmarks
GOZ (Generic Object ZSL Dataset)
The Generix Object Zero-shot Learning (GOZ) dataset is a benchmark dataset for zero-shot learning.
1 paper · 0 benchmarks
GPLA-12 is a new acoustic leakage dataset of gas pipelines involving 12 categories over 684 training/testing acoustic signals.
1 paper · 0 benchmarks
GPR-bench (General‑Purpose Reproducibility Benchmark)
GPR‑bench is an open‑source, multilingual benchmark for regression testing and reproducibility tracking in generative‑AI systems.
1 paper · 0 benchmarks
GPTKB is a large general-domain knowledge base (KB) constructed entirely from a large language model (LLM).
1 paper · 0 benchmarks
The dataset is from google play store applications containing apps from different Google Play Store categories
1 paper · 0 benchmarks
GQN rooms-ring-camera consist of scenes of a variable number of random objects captured in a square room of size 7x7 units.
1 paper · 0 benchmarks
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.