{"url":"/sota/open-knowledge-graph-canonicalization-on-noun","task":{"name":"Open Knowledge Graph Canonicalization","url":"/task/open-knowledge-graph-canonicalization","note":null},"dataset":{"name":"Noun Phrase Canonicalization","url":null},"category":"Graphs","categories":["Graphs","Knowledge Base"],"category_note":null,"description":"Open Information Extraction approaches leads to creation of large Knowledge bases (KB) from the web. The problem with such methods is that their entities and relations are not canonicalized, which leads to storage of redundant and ambiguous facts. For example, an Open KB storing *\\<Barack Obama, was born in, Honolulu\\>* and *\\<Obama, took birth in, Honolulu\\>* doesn't know that *Barack Obama* and *Obama* mean the same entity. Similarly, *took birth in* and *was born in* also refer to the same relation. Problem of Open KB canonicalization involves identifying groups of equivalent entities and relations in the KB.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [CESI: Canonicalizing Open Knowledge Bases using Embeddings and Side Information](https://github.com/malllabiisc/cesi) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Ambiguous dataset","Base Dataset","ReVerb45k"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Ambiguous dataset":null,"Base Dataset":null,"ReVerb45k":null}},"counts":{"rows":1,"rows_with_code":0,"rows_with_paper_page":0,"rows_dated":0,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"Galárraga et al., 2014","metrics":{"Ambiguous dataset":"97.9","Base Dataset":"94.8","ReVerb45k":"98.3"},"uses_additional_data":false,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}