{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-view-approach-to-suggest-moderation","title":"Multi-View Approach to Suggest Moderation Actions in Community Question Answering Sites","arxiv_id":null,"date":"2022-04-04","proceeding":"Information Sciences 2022 4","authors":["Issa Annamoradnejad","Jafar Habibi","Mohammadamin Fazli"],"abstract":"With thousands of new questions posted every day on popular Q&A websites, there is a need for automated and accurate software solutions to replace manual moderation. In this paper, we address the critical drawbacks of crowdsourcing moderation actions in Q&A communities and demonstrate the ability to automate moderation using the latest machine learning models. From a technical point, we propose a multi-view approach that generates three distinct feature groups that examine a question from three different perspectives: 1) question-related features extracted using a BERT-based regression model; 2) context-related features extracted using a named-entity-recognition model; and 3) general lexical features derived using statistical and analytical methods. As a last step, we train a gradient boosting classifier to predict a moderation action. For evaluation purposes, we created a new dataset consisting of 60,000 Stack Overflow questions classified into three choices of moderation actions. Based on cross-validation on the novel dataset, our approach reaches 95.6% accuracy as a multiclass task and outperforms all state-of-the-art and previously-published models. Our results clearly demonstrate the high influence of our feature generation components on the overall success of the classifier.","url_abs":"https://www.sciencedirect.com/science/article/abs/pii/S0020025522003127","url_pdf":"https://www.sciencedirect.com/science/article/abs/pii/S0020025522003127","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"community-question-answering","task_name":"Community Question Answering"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"question-quality-assessment","task_name":"Question Quality Assessment"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[{"slug":"60k-stack-overflow-questions","name":"60k Stack Overflow Questions","full_name":"60k Stack Overflow Questions from 2016-2020 classified into three categories based on their quality"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/question-quality-assessment-on-60k-stack","task":"Question Quality Assessment","dataset":"60k Stack Overflow Questions","model":"Multi-view approach","rank_in_archive_order":1,"of":1,"metrics":{"F1 Score":".917"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}