{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-top-k-keyword-and-top-k-document","title":"Benchmarking Top-K Keyword and Top-K Document Processing with T${}^2$K${}^2$ and T${}^2$K${}^2$D${}^2$","arxiv_id":"1804.07525","date":"2018-04-20","proceeding":null,"authors":["Truica Ciprian-Octavian  UPB","Darmont Jérôme  ERIC","Boicea Alexandru  UPB","Radulescu Florin  UPB"],"abstract":"Top-k keyword and top-k document extraction are very popular text analysis\ntechniques. Top-k keywords and documents are often computed on-the-fly, but\nthey exploit weighted vocabularies that are costly to build. To compare\ncompeting weighting schemes and database implementations, benchmarking is\ncustomary. To the best of our knowledge, no benchmark currently addresses these\nproblems. Hence, in this paper, we present T${}^2$K${}^2$, a top-k keywords and\ndocuments benchmark, and its decision support-oriented evolution\nT${}^2$K${}^2$D${}^2$. Both benchmarks feature a real tweet dataset and queries\nwith various complexities and selectivities. They help evaluate weighting\nschemes and database implementations in terms of computing performance. To\nillustrate our bench-marks' relevance and genericity, we successfully ran\nperformance tests on the TF-IDF and Okapi BM25 weighting schemes, on one hand,\nand on different relational (Oracle, PostgreSQL) and document-oriented\n(MongoDB) database implementations, on the other hand.","url_abs":"http://arxiv.org/abs/1804.07525v1","url_pdf":"http://arxiv.org/pdf/1804.07525v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-top-k-keyword-and-top-k-document","repo_url":"https://github.com/cipriantruica/T2K2D2_Benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}