{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/using-titles-vs-full-text-as-source-for","title":"Using Titles vs. Full-text as Source for Automated Semantic Document Annotation","arxiv_id":"1705.05311","date":"2017-05-15","proceeding":null,"authors":["Lukas Galke","Florian Mai","Alan Schelten","Dennis Brunsch","Ansgar Scherp"],"abstract":"A significant part of the largest Knowledge Graph today, the Linked Open Data\ncloud, consists of metadata about documents such as publications, news reports,\nand other media articles. While the widespread access to the document metadata\nis a tremendous advancement, it is yet not so easy to assign semantic\nannotations and organize the documents along semantic concepts. Providing\nsemantic annotations like concepts in SKOS thesauri is a classical research\ntopic, but typically it is conducted on the full-text of the documents. For the\nfirst time, we offer a systematic comparison of classification approaches to\ninvestigate how far semantic annotations can be conducted using just the\nmetadata of the documents such as titles published as labels on the Linked Open\nData cloud. We compare the classifications obtained from analyzing the\ndocuments' titles with semantic annotations obtained from analyzing the\nfull-text. Apart from the prominent text classification baselines kNN and SVM,\nwe also compare recent techniques of Learning to Rank and neural networks and\nrevisit the traditional methods logistic regression, Rocchio, and Naive Bayes.\nThe results show that across three of our four datasets, the performance of the\nclassifications using only titles reaches over 90% of the quality compared to\nthe classification performance when using the full-text. Thus, conducting\ndocument classification by just using the titles is a reasonable approach for\nautomated semantic annotation and opens up new possibilities for enriching\nKnowledge Graphs.","url_abs":"http://arxiv.org/abs/1705.05311v2","url_pdf":"http://arxiv.org/pdf/1705.05311v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"using-titles-vs-full-text-as-source-for","repo_url":"https://github.com/Quadflor/quadflor","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"document-classification","task_name":"Document Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"knowledge-graphs","task_name":"Knowledge Graphs"},{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[{"method_slug":"svm","method_name":"SVM"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}