{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/kamel-knowledge-analysis-with-multitoken","title":"KAMEL : Knowledge Analysis with Multitoken Entities in Language Models","arxiv_id":null,"date":"2022-11-01","proceeding":"Automated Knowledge Base Construction 2022 11","authors":["Jan-Christoph Kalo","Leandra Fichtel"],"abstract":"Large language models (LMs) have been shown to capture large amounts of relational\r\nknowledge from the pre-training corpus. These models can be probed for this factual knowledge by using cloze-style prompts as demonstrated on the LAMA benchmark. However,\r\nrecent studies have uncovered that results only perform well, because the models are good\r\nat performing educated guesses or recalling facts from the training data. We present a novel\r\nWikidata-based benchmark dataset, KAMEL , for probing relational knowledge in LMs.\r\nIn contrast to previous datasets, it covers a broader range of knowledge, probes for single-,\r\nand multi-token entities, and contains facts with literal values. Furthermore, the evaluation\r\nprocedure is more accurate, since the dataset contains alternative entity labels and deals\r\nwith higher-cardinality relations. Instead of performing the evaluation on masked language\r\nmodels, we present results for a variety of recent causal LMs in a few-shot setting. We show\r\nthat indeed novel models perform very well on LAMA, achieving a promising F1-score of\r\n52.90%, while only achieving 17.62% on KAMEL. Our analysis shows that even large language models are far from being able to memorize all varieties of relational knowledge that\r\nis usually stored knowledge graphs.","url_abs":"https://www.akbc.ws/2022/assets/pdfs/15_kamel_knowledge_analysis_with_.pdf","url_pdf":"https://www.akbc.ws/2022/assets/pdfs/15_kamel_knowledge_analysis_with_.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"kamel-knowledge-analysis-with-multitoken","repo_url":"https://github.com/JanKalo/KAMEL","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"knowledge-graphs","task_name":"Knowledge Graphs"},{"task_slug":"probing-language-models","task_name":"Probing Language Models"}],"methods":[{"method_slug":"lama","method_name":"LAMA"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[{"slug":"kamel","name":"KAMEL","full_name":"Knowledge Analysis with Multitoken Entities in Language Models"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/probing-language-models-on-kamel","task":"Probing Language Models","dataset":"KAMEL","model":"OPT-13b","rank_in_archive_order":1,"of":1,"metrics":{"Average F1":"17.62"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}