{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-off-topic-memento-toolkit","title":"The Off-Topic Memento Toolkit","arxiv_id":"1806.06870","date":"2018-06-18","proceeding":null,"authors":["Shawn M. Jones","Michele C. Weigle","Michael L. Nelson"],"abstract":"Web archive collections are created with a particular purpose in mind. A\ncurator selects seeds, or original resources, which are then captured by an\narchiving system and stored as archived web pages, or mementos. The systems\nthat build web archive collections are often configured to revisit the same\noriginal resource multiple times. This is incredibly useful for understanding\nan unfolding news story or the evolution of an organization. Unfortunately,\nover time, some of these original resources can go off-topic and no longer suit\nthe purpose for which the collection was originally created. They can go\noff-topic due to web site redesigns, changes in domain ownership, financial\nissues, hacking, technical problems, or because their content has moved on from\nthe original topic. Even though they are off-topic, the archiving system will\nstill capture them, thus it becomes imperative to anyone performing research on\nthese collections to identify these off-topic mementos. Hence, we present the\nOff-Topic Memento Toolkit, which allows users to detect off-topic mementos\nwithin web archive collections. The mementos identified by this toolkit can\nthen be separately removed from a collection or merely excluded from downstream\nanalysis. The following similarity measures are available: byte count, word\ncount, cosine similarity, Jaccard distance, S{\\o}rensen-Dice distance, Simhash\nusing raw text content, Simhash using term frequency, and Latent Semantic\nIndexing via the gensim library. We document the implementation of each of\nthese similarity measures. We possess a gold standard dataset generated by\nmanual analysis, which contains both off-topic and on-topic mementos. Using\nthis gold standard dataset, we establish a default threshold corresponding to\nthe best F1 score for each measure. We also provide an overview of potential\nfuture directions that the toolkit may take.","url_abs":"http://arxiv.org/abs/1806.06870v1","url_pdf":"http://arxiv.org/pdf/1806.06870v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-off-topic-memento-toolkit","repo_url":"https://github.com/oduwsdl/off-topic-memento-toolkit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}