Home › Datasets › task › Knowledge Graphs

Knowledge Graphs datasets

archive 2025-07-28

46 datasets carry the task tag "Knowledge Graphs" (the task itself: Knowledge Graphs), ordered by the archive's paper count. Page 1 of 1: 46 shown of 46. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Knowledge Graphs datasets 1–46 of 46

ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges.
852 papers · 1 benchmark
FB15k (Freebase 15K)
The FB15k dataset contains knowledge base relation triples and textual mentions of Freebase entity pairs.
641 papers · 6 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other.
105 papers · 2 benchmarks
MetaQA (MoviE Text Audio QA)
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
The Semantic Scholar corpus (S2) is composed of titles from scientific papers published in machine learning conferences and journals from 1985 to 2017, split by year (33 timesteps).
75 papers · 0 benchmarks
DBP15k contains four language-specific KGs that are respectively extracted from English (En), Chinese (Zh), French (Fr) and Japanese (Ja) DBpedia, each of which contains around 65k-106k entities.
65 papers · 3 benchmarks
ComplexWebQuestions is a dataset for answering complex questions that require reasoning over multiple web snippets.
62 papers · 2 benchmarks
OpenDialKG contains utterance from 15K human-to-human role-playing dialogs is manually annotated with ground-truth reference to corresponding entities and paths from a large-scale KG with 1M+ facts.
55 papers · 0 benchmarks
Jericho is a learning environment for man-made Interactive Fiction (IF) games.
54 papers · 0 benchmarks
MMKG is a collection of three knowledge graphs for link prediction and entity matching research.
49 papers · 3 benchmarks
Contains around 200K dialogs with a total of 1.6M turns.
40 papers · 0 benchmarks
OGB-LSC (OGB Large-Scale Challenge)
OGB Large-Scale Challenge (OGB-LSC) is a collection of three real-world datasets for advancing the state-of-the-art in large-scale graph ML.
34 papers · 3 benchmarks
Holl-E is a dataset containing movie chats wherein each response is explicitly generated by copying and/or modifying sentences from unstructured background knowledge such as plots, comments and reviews about the movie.
30 papers · 0 benchmarks
RoboCup is an initiative in which research groups compete by enabling their robots to play football matches.
20 papers · 0 benchmarks
One of the largest commonsense knowledge bases available, describing over 2 million disambiguated concepts and activities, connected by over 18 million assertions.
20 papers · 0 benchmarks
A large-scale aerial farmland image dataset for semantic segmentation of agricultural patterns.
18 papers · 0 benchmarks
This relational database consists of 24 unique names in two families (they have equivalent structures).
10 papers · 0 benchmarks
A corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities.
10 papers · 0 benchmarks
NLPContributionGraph was introduced as Task 11 at SemEval 2021 for the first time.
8 papers · 0 benchmarks
ForecastQA is a question-answering dataset consisting of 10,392 event forecasting questions, which have been collected and verified via crowdsourcing efforts.
7 papers · 0 benchmarks
JerichoWorld is a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives.
6 papers · 2 benchmarks
KnowledgeNet is a benchmark dataset for the task of automatically populating a knowledge base (Wikidata) with facts expressed in natural language text on the web.
6 papers · 0 benchmarks
SemOpenAlex is an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
KG20C (A scholarly knowledge graph benchmark dataset)
KG20C is a Knowledge Graph about high quality papers from 20 top computer science Conferences.
5 papers · 1 benchmark
A small RDF Knowledge Graph using FOAF and VCard.
5 papers · 0 benchmarks
The WorldKG knowledge graph is a comprehensive large-scale geospatial knowledge graph based on OpenStreetMap that provides a semantic representation of geographic entities from over 188 countries.
5 papers · 0 benchmarks
The FB15k-237-low dataset is a variation of the FB15k-237 dataset where relations with a low number of triplets are kept.
3 papers · 0 benchmarks
WikiWiki is a dataset for understanding entities and their place in a taxonomy of knowledge—their types.
3 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
The dataset is constructed from an Amazon review corpus by integrating both user-agent dialogue and custom knowledge graphs for recommendation.
2 papers · 0 benchmarks
A new benchmark dataset for simple question answering over knowledge graphs that was created by mapping SimpleQuestions entities and predicates from Freebase to DBpedia.
2 papers · 0 benchmarks
AISECKG (AISecKG: Knowledge Graph Dataset for Cybersecurity Education)
Cybersecurity education is exceptionally challenging as it involves learning the complex attacks; tools and developing critical problem-solving skills to defend the systems.
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
DTBM (Digital Twin Benchmark Model)
DTBM is a benchmark dataset for Digital Twins that reflects these characteristics and look into the scaling challenges of different knowledge graph technologies.
1 paper · 0 benchmarks
ENT-DESC involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter…
1 paper · 1 benchmark
Contains 1000 semantic queries and the corresponding English, German and Portuguese verbalizations for EventKG - an event-centric knowledge graph with more than 970 thousand events.
1 paper · 0 benchmarks
The FB1.5M dataset is a benchmark for Knowledge Graph Completion.
1 paper · 0 benchmarks
The KACC benchmark consists of three subtasks that can be applied to knowledge graphs: knowledge abstraction, knowledge concretization and knowledge completion.
1 paper · 0 benchmarks
Analogical reasoning is fundamental to human cognition and holds an important place in various fields.
1 paper · 1 benchmark
A preliminary dataset of related tables and a corresponding set of natural language questions.
1 paper · 0 benchmarks
TextWorld KG is a dynamic Knowledge Graph (KG) extraction dataset.
1 paper · 0 benchmarks
Dataset Description: Summarized Wiki Articles with TTL Knowledge Graphs Overview This dataset comprises 500 summarized Wikipedia articles, each accompanied by a corresponding TTL knowledge graph.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.