{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mega-cov-a-billion-scale-dataset-of-65","title":"Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19","arxiv_id":"2005.06012","date":"2020-05-02","proceeding":"EACL 2021 2","authors":["Muhammad Abdul-Mageed","AbdelRahim Elmadany","El Moatez Billah Nagoudi","Dinesh Pabbi","Kunal Verma","Rannie Lin"],"abstract":"We describe Mega-COV, a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 268 countries), longitudinal (goes as back as 2007), multilingual (comes in 100+ languages), and has a significant number of location-tagged tweets (~169M tweets). We release tweet IDs from the dataset. We also develop and release two powerful models, one for identifying whether or not a tweet is related to the pandemic (best F1=97%) and another for detecting misinformation about COVID-19 (best F1=92%). A human annotation study reveals the utility of our models on a subset of Mega-COV. Our data and models can be useful for studying a wide host of phenomena related to the pandemic. Mega-COV and our models are publicly available.","url_abs":"https://arxiv.org/abs/2005.06012v4","url_pdf":"https://arxiv.org/pdf/2005.06012v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mega-cov-a-billion-scale-dataset-of-65","repo_url":"https://github.com/UBC-NLP/megacov","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"misinformation","task_name":"Misinformation"}],"methods":[],"datasets_introduced":[{"slug":"mega-cov","name":"Mega-COV","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}