{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-mafiascum-dataset-a-large-text-corpus-for","title":"The Mafiascum Dataset: A Large Text Corpus for Deception Detection","arxiv_id":"1811.07851","date":"2018-11-19","proceeding":null,"authors":["Bob de Ruiter","George Kachergis"],"abstract":"Detecting deception in natural language has a wide variety of applications, but because of its hidden nature there are currently no public, large-scale sources of labeled deceptive text. This work introduces the Mafiascum dataset [1], a collection of over 700 games of Mafia, in which players are randomly assigned either deceptive or non-deceptive roles and then interact via forum postings. Over 9000 documents were compiled from the dataset, which each contained all messages written by a single player in a single game. This corpus was used to construct a set of hand-picked linguistic features based on prior deception research, as well as a set of average word vectors enriched with subword information. A logistic regression classifier fit on a combination of these feature sets achieved an average precision of 0.39 (chance = 0.26) and an AUROC of 0.68 on 5000+ word documents. On 50+ word documents, an average precision of 0.29 (chance = 0.23) and an AUROC of 0.59 was achieved. [1] https://bitbucket.org/bopjesvla/thesis/src","url_abs":"https://arxiv.org/abs/1811.07851v3","url_pdf":"https://arxiv.org/pdf/1811.07851v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-mafiascum-dataset-a-large-text-corpus-for","repo_url":"https://bitbucket.org/bopjesvla/thesis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"deception-detection","task_name":"Deception Detection"}],"methods":[{"method_slug":"logistic-regression","method_name":"Logistic Regression"}],"datasets_introduced":[{"slug":"mafiascum","name":"Mafiascum","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.07851","atlas_url":"https://app.syntology.ai/?focus=1811.07851","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}