Home › Datasets › language › Chinese
Chinese datasets
archive 2025-07-28
460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 1 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Chinese datasets 1–48 of 460
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
NeRF (Neural Radiance Fields)
Neural Radiance Fields (NeRF) is a method for synthesizing novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views.
3,892 papers · 1 benchmark
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
CelebA (CelebFaces Attributes Dataset)
CelebFaces Attributes dataset contains 202,599 face images of the size 178×218 from 10,177 celebrities, each annotated with 40 binary labels indicating facial attributes like hair color, gender and age.
3,477 papers · 17 benchmarks
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
ShapeNet is a large scale repository for 3D CAD models developed by researchers from Stanford University, Princeton University and the Toyota Technological Institute at Chicago, USA.
1,947 papers · 13 benchmarks
GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers.
1,881 papers · 7 benchmarks
The ModelNet40 dataset contains synthetic object point clouds.
1,406 papers · 15 benchmarks
mini-Imagenet is proposed by Matching Networks for One Shot Learning .
1,345 papers · 21 benchmarks
The ADE20K semantic segmentation dataset contains more than 20K scene-centric images exhaustively annotated with pixel-level objects and object parts labels.
1,213 papers · 32 benchmarks
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips.
744 papers · 5 benchmarks
DIV2K is a popular single-image super-resolution dataset which contains 1,000 images with different scenes and is splitted to 800 for training, 100 for validation and 100 for testing.
654 papers · 3 benchmarks
ImageNet-C is an open source data set that consists of algorithmically generated corruptions (blur, noise) applied to the ImageNet test-set.
602 papers · 4 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
FB15k-237 is a link prediction dataset created from FB15k.
452 papers · 3 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
DailyDialog is a high-quality multi-turn open-domain English dialog dataset.
399 papers · 2 benchmarks
Netflix Prize consists of about 100,000,000 ratings for 17,770 movies given by 480,189 users.
370 papers · 1 benchmark
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics.
348 papers · 5 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
ETT (Electricity Transformer Temperature)
The Electricity Transformer Temperature (ETT) is a crucial indicator in the electric power long-term deployment.
321 papers · 20 benchmarks
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
MSMT17 (Multi Scene Multi Time dataset for person re-id)
MSMT17 is a multi-scene multi-time person re-identification dataset.
275 papers · 6 benchmarks
Description: 10,000 People - Human Pose Recognition Data.
265 papers · 1 benchmark
BSDS500 (Berkeley Segmentation Dataset 500)
Berkeley Segmentation Data Set 500 (BSDS500) is a standard benchmark for contour detection.
261 papers · 8 benchmarks
The NCI1 dataset comes from the cheminformatics domain, where each input graph is used as representation of a chemical compound: each vertex stands for an atom of the molecule, and edges between vertices represent bonds between atoms.
260 papers · 2 benchmarks
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
ImageNet Long-Tailed is a subset of /dataset/imagenet dataset consisting of 115.8K images from 1000 categories, with maximally 1280 images per class and minimally 5 images per class.
219 papers · 3 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9.
178 papers · 3 benchmarks
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
RAF-DB (Real-world Affective Faces)
The Real-world Affective Faces Database (RAF-DB) is a dataset for facial expression.
172 papers · 3 benchmarks
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
Total-Text is a text detection dataset that consists of 1,555 images with a variety of text types including horizontal, multi-oriented, and curved text instances.
156 papers · 2 benchmarks
The MSRA-TD500 dataset is a text detection dataset that contains 300 training images and 200 test images.
124 papers · 1 benchmark
The Microsoft Academic Graph is a heterogeneous graph containing scientific publication records, citation relationships between those publications, as well as authors, institutions, journals, conferences, and fields of study.
124 papers · 0 benchmarks
JFT-300M is an internal Google dataset used for training image classification models.
123 papers · 1 benchmark
GTEA (Georgia Tech Egocentric Activity)
The Georgia Tech Egocentric Activities (GTEA) dataset contains seven types of daily activities such as making sandwich, tea, or coffee.
120 papers · 2 benchmarks
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.