{"url":"/task/open-knowledge-graph-canonicalization","name":"Open Knowledge Graph Canonicalization","slug":"open-knowledge-graph-canonicalization","description_markdown":"Open Information Extraction approaches leads to creation of large Knowledge bases (KB) from the web. The problem with such methods is that their entities and relations are not canonicalized, which leads to storage of redundant and ambiguous facts. For example, an Open KB storing *\\<Barack Obama, was born in, Honolulu\\>* and *\\<Obama, took birth in, Honolulu\\>* doesn't know that *Barack Obama* and *Obama* mean the same entity. Similarly, *took birth in* and *was born in* also refer to the same relation. Problem of Open KB canonicalization involves identifying groups of equivalent entities and relations in the KB.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [CESI: Canonicalizing Open Knowledge Bases using Embeddings and Side Information](https://github.com/malllabiisc/cesi) )</span>","categories":[{"name":"Graphs","url":"/area/graphs"},{"name":"Knowledge Base","url":"/area/knowledge-base"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":3,"papers_with_code":2,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/open-knowledge-graph-canonicalization-on-noun","slug":"open-knowledge-graph-canonicalization-on-noun","dataset":"Noun Phrase Canonicalization","dataset_url":null,"rows_in_archive":1,"metrics":["Ambiguous dataset","Base Dataset","ReVerb45k"],"first_row_in_archive_order":{"model":"Galárraga et al., 2014","paper_title":null,"paper_url":null,"paper_date":"","arxiv_id":null,"code_links":[],"syntology":null}}],"datasets":[{"url":"/dataset/reverb45k","name":"ReVerb45K","full_name":"ReVerb45K","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/knowledge-graphs","name":"Knowledge Graphs"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":2,"of":2,"tagged_in_all":3,"items":[{"url":"/paper/combo-a-complete-benchmark-for-open-kg","title":"COMBO: A Complete Benchmark for Open KG Canonicalization","date":"2023-02-08","arxiv_id":"2302.03905","repositories_listed":1,"syntology":null},{"url":"/paper/cesi-canonicalizing-open-knowledge-bases","title":"CESI: Canonicalizing Open Knowledge Bases using Embeddings and Side Information","date":"2019-02-01","arxiv_id":"1902.00172","repositories_listed":1,"syntology":null}],"syntology_records":0,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}