Browse State-of-the-Art › Knowledge Base Population
Knowledge Base Population
32 papers with code · 1 benchmark · 3 datasets archive 2025-07-28
Knowledge base population is the task of filling the incomplete elements of a given knowledge base by automatically processing a large corpus of text.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LM-KBC 2023 (1 row) | VE-BERT | Expanding the Vocabulary of BERT for Knowledge Base Construction | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 32 papers with code (140 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Sep 2021 2 repositories listedExperimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task.
-
17 Apr 2021 2 repositories listedRecently, there has been a promising direction in evaluating language models in the same way we would evaluate knowledge bases, and the task of slot filling is the most suitable to this intent.
-
12 Jun 2020 2 repositories listedClinical trials predicate subject eligibility on a diversity of criteria ranging from patient demographics to food allergies.
-
1 Sep 2017 2 repositories listedThe combination of better supervised data and a more appropriate high-capacity model enables much better relation extraction performance.
-
2 Jun 2017 2 repositories listedIf not, what characteristics of a dataset determine the performance of MF and TF models?
-
29 Jan 2024 1 repository listedKnowledge base population seeks to expand knowledge graphs with facts that are typically extracted from a text corpus.
-
12 Oct 2023 1 repository listedTo address this, we present Vocabulary Expandable BERT for knowledge base construction, which expand the language model's vocabulary while preserving semantic embeddings for newly added words.
-
20 Apr 2023 1 repository listedWe show that CKBP v2 serves as a challenging and representative evaluation dataset for the CSKB Population task, while its development set aids in selecting a population model that leads to improved knowledge…
-
14 Oct 2022 1 repository listedWe propose PseudoReasoner, a semi-supervised learning framework for CSKB population that uses a teacher model pre-trained on CSKBs to provide pseudo labels on the unlabeled candidate dataset for a student model to learn…
-
1 Jul 2022 1 repository listedEvents are inter-related in documents.
-
30 May 2022 1 repository listedEvents are inter-related in documents.
-
1 May 2022 1 repository listedAutomatic extraction of event structures from text is a promising way to extract important facts from the evergrowing amount of biomedical literature.
-
10 Jan 2022 1 repository listedWe present an open-source and extensible knowledge extraction toolkit DeepKE, supporting complicated low-resource, document-level and multimodal scenarios in the knowledge base population.
-
5 Apr 2021 1 repository listedRelation extraction from text is an important task for automatic knowledge base population.
-
1 Nov 2020 1 repository listedBiomedical event extraction from natural text is a challenging task as it searches for complex and often nested structures describing specific relationships between multiple molecular entities, such as genes, proteins,…
-
2 May 2020 1 repository listedMost existing methods train with a small number of negative samples for each positive instance in these datasets to save computational costs.
-
1 Nov 2019 1 repository listedKnowledgeNet is a benchmark dataset for the task of automatically populating a knowledge base (Wikidata) with facts expressed in natural language text on the web.
-
1 Jul 2019 1 repository listedUnderstanding the structures of political debates (which actors make what claims) is essential for understanding democratic political decision making.
-
1 Aug 2018 1 repository listedIn this paper, we come up with a feature adaptation approach for cross-lingual relation classification, which employs a generative adversarial network (GAN) to transfer feature representations from one language with…
-
1 Aug 2018 1 repository listedIn this work, we propose a model which alleviates the need for such disambiguators by jointly learning NER and MD taggers in languages for which one can provide a list of candidate morphological analyses.
-
1 Aug 2018 1 repository listedWe introduce INCEpTION, a new annotation platform for tasks including interactive and semantic annotation (e.
-
1 Jul 2018 1 repository listedState-of-the-art relation extraction approaches are only able to recognize relationships between mentions of entity arguments stated explicitly in the text and typically localized to the same sentence.
-
1 Jul 2018 1 repository listedState-of-the-art knowledge base completion (KBC) models predict a score for every known or unknown fact via a latent factorization over entity and relation embeddings.
-
3 Jun 2018 1 repository listedKnowledge Base Population (KBP) is the task of building or extending a knowledge base from text, and systems for KBP have grown in capability and scope.
-
1 May 2018 1 repository listed
-
1 May 2018 1 repository listed
-
1 May 2018 1 repository listed
-
10 Apr 2017 1 repository listedWe describe the SemEval task of extracting keyphrases and relations between them from scientific documents, which is crucial for understanding which publications describe which processes, tasks and materials.
-
TweeTime : A Minimally Supervised Method for Recognizing and Normalizing Time Expressions in Twitter1 Nov 2016 1 repository listed
-
1 Jun 2016 1 repository listed
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections