{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/structural-regularities-in-text-based-entity","title":"Structural Regularities in Text-based Entity Vector Spaces","arxiv_id":"1707.07930","date":"2017-07-25","proceeding":null,"authors":["Christophe Van Gysel","Maarten de Rijke","Evangelos Kanoulas"],"abstract":"Entity retrieval is the task of finding entities such as people or products\nin response to a query, based solely on the textual documents they are\nassociated with. Recent semantic entity retrieval algorithms represent queries\nand experts in finite-dimensional vector spaces, where both are constructed\nfrom text sequences.\n  We investigate entity vector spaces and the degree to which they capture\nstructural regularities. Such vector spaces are constructed in an unsupervised\nmanner without explicit information about structural aspects. For concreteness,\nwe address these questions for a specific type of entity: experts in the\ncontext of expert finding. We discover how clusterings of experts correspond to\ncommittees in organizations, the ability of expert representations to encode\nthe co-author graph, and the degree to which they encode academic rank. We\ncompare latent, continuous representations created using methods based on\ndistributional semantics (LSI), topic models (LDA) and neural networks\n(word2vec, doc2vec, SERT). Vector spaces created using neural methods, such as\ndoc2vec and SERT, systematically perform better at clustering than LSI, LDA and\nword2vec. When it comes to encoding entity relations, SERT performs best.","url_abs":"http://arxiv.org/abs/1707.07930v1","url_pdf":"http://arxiv.org/pdf/1707.07930v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"structural-regularities-in-text-based-entity","repo_url":"https://github.com/cvangysel/SERT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"entity-retrieval","task_name":"Entity Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[{"method_slug":"lda","method_name":"LDA"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}