Home › Datasets › language › Indonesian
Indonesian datasets
archive 2025-07-28
44 datasets carry the language tag "Indonesian", ordered by the archive's paper count. Page 1 of 1: 44 shown of 44. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Indonesian datasets 1–44 of 44
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
AFW (Annotated Faces in the Wild)
AFW (Annotated Faces in the Wild) is a face detection dataset that contains 205 images with 468 faces.
154 papers · 1 benchmark
The Microsoft Academic Graph is a heterogeneous graph containing scientific publication records, citation relationships between those publications, as well as authors, institutions, journals, conferences, and fields of study.
124 papers · 0 benchmarks
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
CORD (Consolidated Receipt Dataset for Post-OCR Parsing)
OCR is inevitably linked to NLP since its final output is in text.
100 papers · 1 benchmark
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages.
26 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks
X-FACT is a large publicly available multilingual dataset for factual verification of naturally existing real-world claims.
16 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
A large-scale Indonesian summarization dataset consisting of harvested articles from Liputan6.com, an online news portal, resulting in 215,827 document-summary pairs.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
IndoNLG is a benchmark to measure natural language generation (NLG) progress in three low-resource—yet widely spoken—languages of Indonesia: Indonesian, Javanese, and Sundanese.
4 papers · 0 benchmarks
IndoNLI is the first human-elicited NLI dataset for Indonesian consisting of nearly 18K sentence pairs annotated by crowd workers and experts.
4 papers · 0 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
QASiNa (Question Answering Sirah Nabawiyah)
Question Answering Sirah Nabawiyah (QASiNa) Dataset is a reading comprehension dataset consists of QA from Sirah Nabawiyah literature in Indonesian Language
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
NERGRIT involves machine learning based NLP Tools and a corpus used for Indonesian Named Entity Recognition, Statement Extraction, and Sentiment Analysis.
0 papers · 1 benchmark
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.