Home › Datasets › language › Japanese

Japanese datasets

archive 2025-07-28

111 datasets carry the language tag "Japanese", ordered by the archive's paper count. Page 3 of 3: 15 shown of 111. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Japanese datasets 97–111 of 111

PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
A large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions.
1 paper · 0 benchmarks
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
SimpleStories is a dataset of >2 million model-generated short stories.
1 paper · 0 benchmarks
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
YJMob100K (YJMob100K: City-Scale and Longitudinal Dataset of Anonymized Human Mobility Trajectories)
Modeling and predicting human mobility trajectories in urban areas is an essential task for various applications including transportation modeling, disaster management, and urban planning.
1 paper · 0 benchmarks
jaCappella is a corpus of Japanese a cappella vocal ensembles (jaCappella corpus) for vocal ensemble separation and synthesis.
1 paper · 0 benchmarks
ADFI (Anomaly Detection Datasets for Visual Inspection)
ADFI Dataset is an image dataset for anomaly detection methods with a focus on industrial inspection.
0 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
0 papers · 0 benchmarks
JTES (Japanese Twitter-based Emotional Speech)
We designed an emotional speech database that can be used for emotion recognition as well as recognition and synthesis of speech with various emotions.
0 papers · 0 benchmarks
Robbie Williams is a dataset of 65 songs by Robbie Williams.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.