Home › Datasets › task › Word Embeddings

Word Embeddings datasets

archive 2025-07-28

52 datasets carry the task tag "Word Embeddings" (the task itself: Word Embeddings), ordered by the archive's paper count. Page 1 of 2: 48 shown of 52. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Word Embeddings datasets 1–48 of 52

ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges.
852 papers · 1 benchmark
FrameNet is a linguistic knowledge graph containing information about lexical and predicate argument semantics of the English language.
444 papers · 0 benchmarks
BookCorpus is a large collection of free novel books written by unpublished authors, which contains 11,038 books (around 74M sentences and 1G words) of 16 different sub-genres (e.g., Romance, Historical, Adventure, etc.).
344 papers · 1 benchmark
WiC (Words in Context)
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
BioASQ (Biomedical Semantic Indexing and Question Answering)
BioASQ is a question answering dataset.
192 papers · 1 benchmark
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
WinoBias contains 3,160 sentences, split equally for development and test, created by researchers familiar with the project.
134 papers · 0 benchmarks
PadChest is a labeled large-scale, high resolution chest x-ray dataset for the automated exploration of medical images along with their associated reports.
116 papers · 0 benchmarks
WikiMatrix is a dataset of parallel sentences in the textual content of Wikipedia for all possible language pairs.
91 papers · 0 benchmarks
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
CELEX database comprises three different searchable lexical databases, Dutch, English and German.
59 papers · 0 benchmarks
ISEAR (International Survey on Emotion Antecedents and Reactions)
Over a period of many years during the 1990s, a large group of psychologists all over the world collected data in the ISEAR project, directed by Klaus R.
55 papers · 0 benchmarks
PanLex translates words in thousands of languages.
36 papers · 0 benchmarks
EVALution dataset is evenly distributed among the three classes (hypernyms, co-hyponyms and random) and involves three types of parts of speech (noun, verb, adjective).
28 papers · 0 benchmarks
There are now many computer programs for automatically determining the sense of a word in context (Word Sense Disambiguation or WSD).
19 papers · 0 benchmarks
The first parallel corpus composed from United Nations documents published by the original data creator.
18 papers · 0 benchmarks
An expert-annotated word similarity dataset which provides a highly reliable, yet challenging, benchmark for rare word representation techniques.
10 papers · 0 benchmarks
Polyglot-NER builds massive multilingual annotators with minimal human expertise and intervention.
10 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
ERA (Event Recognition in Aerial videos)
Consists of 2,864 videos each with a label from 25 different classes corresponding to an event unfolding 5 seconds.
8 papers · 0 benchmarks
SEND (Stanford Emotional Narratives Dataset)
SEND (Stanford Emotional Narratives Dataset) is a set of rich, multimodal videos of self-paced, unscripted emotional narratives, annotated for emotional valence over time.
8 papers · 0 benchmarks
A new challenging dataset that can be used for many pattern recognition tasks.
7 papers · 1 benchmark
Introduces three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis, respectively.
6 papers · 0 benchmarks
SemEval 2014 is a collection of datasets used for the Semantic Evaluation (SemEval) workshop, an annual event that focuses on the evaluation and comparison of systems that can analyze diverse semantic phenomena in text.
6 papers · 0 benchmarks
WNLaMPro (WordNet Language Model Probing)
The WordNet Language Model Probing (WNLaMPro) dataset consists of relations between keywords and words.
6 papers · 0 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
A large-scale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.
5 papers · 0 benchmarks
In this paper, we present AnlamVer, which is a semantic model evaluation dataset for Turkish designed to evaluate word similarity and word relatedness tasks while discriminating those two relations from each other.
4 papers · 0 benchmarks
Some Like it Hoax is a fake news detection dataset consisting of 15,500 Facebook posts and 909,236 users.
4 papers · 0 benchmarks
The WikiSem500 dataset contains around 500 per-language cluster groups for English, Spanish, German, Chinese, and Japanese (a total of 13,314 test cases).
4 papers · 0 benchmarks
DICE: a Dataset of Italian Crime Event news (from Gazzetta di Modena [2011-2021])
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
The dataset consists of the features associated with 402 5-second sound samples.
3 papers · 0 benchmarks
The IndicNLP corpus is a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families.
3 papers · 0 benchmarks
Contains more than 6500 words semantically grouped under 110 categories.
3 papers · 0 benchmarks
Chinese Gigaword corpus consists of 2.2M of headline-document pairs of news stories covering over 284 months from two Chinese newspapers, namely the Xinhua News Agency of China (XIN) and the Central News Agency of Taiwan (CNA).
2 papers · 0 benchmarks
The Climate Change Claims dataset for generating fact checking summaries contains claims broadly related to climate change and global warming from climatefeedback.org.
2 papers · 0 benchmarks
This dataset contains information about Japanese word similarity including rare words.
2 papers · 0 benchmarks
The MUSE dataset contains bilingual dictionaries for 110 pairs of languages.
2 papers · 2 benchmarks
SART is a collection of three datasets for Similarity, Analogies and Relatedness for the Tatar language.
2 papers · 0 benchmarks
Collects a huge number of job descriptions from Dice.com - one of the most popular career website about Tech jobs in USA.
2 papers · 0 benchmarks
The pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
2 papers · 0 benchmarks
We provide a Mikolov-style word-analogy evaluation set specifically for Bangla, with a sample size of 16678, as well as a translated and curated version of the Mikolov dataset, which contains 10594 samples for cross-lingual research.
1 paper · 0 benchmarks
A distant supervision dataset by linking the entire English ClueWeb09 corpus to Freebase.
1 paper · 0 benchmarks
A dataset of sentence pairs annotated following the formalization.
1 paper · 0 benchmarks
A new word analogy task dataset for Indonesian.
1 paper · 0 benchmarks
LARQS (An Evaluation Dataset for Chinese Codex Word Embedding Model)
Word embedding is a modern distributed word representations approach widely used in many natural language processing tasks.
1 paper · 0 benchmarks
McQueen dataset contains 15k visual conversations and over 80k queries where each one is associated with a fully-specified rewrite version.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.