{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/twitter-as-a-lifeline-human-annotated-twitter","title":"Twitter as a Lifeline: Human-annotated Twitter Corpora for NLP of Crisis-related Messages","arxiv_id":"1605.05894","date":"2016-05-19","proceeding":"LREC 2016 5","authors":["Muhammad Imran","Prasenjit Mitra","Carlos Castillo"],"abstract":"Microblogging platforms such as Twitter provide active communication channels\nduring mass convergence and emergency events such as earthquakes, typhoons.\nDuring the sudden onset of a crisis situation, affected people post useful\ninformation on Twitter that can be used for situational awareness and other\nhumanitarian disaster response efforts, if processed timely and effectively.\nProcessing social media information pose multiple challenges such as parsing\nnoisy, brief and informal messages, learning information categories from the\nincoming stream of messages and classifying them into different classes among\nothers. One of the basic necessities of many of these tasks is the availability\nof data, in particular human-annotated data. In this paper, we present\nhuman-annotated Twitter corpora collected during 19 different crises that took\nplace between 2013 and 2015. To demonstrate the utility of the annotations, we\ntrain machine learning classifiers. Moreover, we publish first largest word2vec\nword embeddings trained on 52 million crisis-related tweets. To deal with\ntweets language issues, we present human-annotated normalized lexical resources\nfor different lexical variations.","url_abs":"http://arxiv.org/abs/1605.05894v2","url_pdf":"http://arxiv.org/pdf/1605.05894v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"twitter-as-a-lifeline-human-annotated-twitter","repo_url":"https://github.com/konstapo/2022-fake-news-mediaeval-task","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"disaster-response","task_name":"Disaster Response"},{"task_slug":"humanitarian","task_name":"Humanitarian"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1605.05894","atlas_url":"https://app.syntology.ai/?focus=1605.05894","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}