{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cqasumm-building-references-for-community","title":"CQASUMM: Building References for Community Question Answering Summarization Corpora","arxiv_id":"1811.04884","date":"2018-11-12","proceeding":null,"authors":["Tanya Chowdhury","Tanmoy Chakraborty"],"abstract":"Community Question Answering forums such as Quora, Stackoverflow are rich\nknowledge resources, often catering to information on topics overlooked by\nmajor search engines. Answers submitted to these forums are often elaborated,\ncontain spam, are marred by slurs and business promotions. It is difficult for\na reader to go through numerous such answers to gauge community opinion. As a\nresult summarization becomes a prioritized task for CQA forums. While a number\nof efforts have been made to summarize factoid CQA, little work exists in\nsummarizing non-factoid CQA. We believe this is due to the lack of a\nconsiderably large, annotated dataset for CQA summarization. We create CQASUMM,\nthe first huge annotated CQA summarization dataset by filtering the 4.4 million\nYahoo! Answers L6 dataset. We sample threads where the best answer can double\nup as a reference summary and build hundred word summaries from them. We treat\nother answers as candidates documents for summarization. We provide a script to\ngenerate the dataset and introduce the new task of Community Question Answering\nSummarization. Multi document summarization has been widely studied with news\narticle datasets, especially in the DUC and TAC challenges using news corpora.\nHowever documents in CQA have higher variance, contradicting opinion and lesser\namount of overlap. We compare the popular multi document summarization\ntechniques and evaluate their performance on our CQA corpora. We look into the\nstate-of-the-art and understand the cases where existing multi document\nsummarizers (MDS) fail. We find that most MDS workflows are built for the\nentirely factual news corpora, whereas our corpus has a fair share of opinion\nbased instances too. We therefore introduce OpinioSumm, a new MDS which\noutperforms the best baseline by 4.6% w.r.t ROUGE-1 score.","url_abs":"http://arxiv.org/abs/1811.04884v1","url_pdf":"http://arxiv.org/pdf/1811.04884v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cqasumm-building-references-for-community","repo_url":"https://bitbucket.org/tanya14109/cqasumm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"community-question-answering","task_name":"Community Question Answering"},{"task_slug":"document-summarization","task_name":"Document Summarization"},{"task_slug":"multi-document-summarization","task_name":"Multi-Document Summarization"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[{"slug":"cqasumm","name":"CQASUMM","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.04884","atlas_url":"https://app.syntology.ai/?focus=1811.04884","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}