{"url":"/dataset/echo-corpus","name":"Echo Corpus","full_name":null,"description_markdown":"A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner:\r\n1. **Base Corpus** We have obtained access to a random sample of 10% of all public tweets posted in May and June 2016 – the peak use of the echo.\r\n2. **Raw Echo Corpus** Searching the base corpus, we extracted all tweets containing the echo symbol, resulting in 803,539 tweets posted by 418,624 users. Filtering out non-English Tweets and users who used the echo less than three times we were left with ∼7K users. \r\n3. **Echo Corpus** We used Twitter API to obtain the most recent tweets (up to 3.2K) of each of the users remainingin the English list. This process resulted in ∼18M tweets posted by 7,073 users. Some of the accounts we found using the echo were already suspended or deleted at the\r\ntime of collection, thus their tweets were not retrievable.\r\nRelevant footnotes:\r\n- The echo is found in tweets written in multiple languages, particularly in East-Asian languages of which the user based is known for heavy use of ascii art and kaomoji (McCulloch 2019).\r\n- The data was collected in December 2016, amidst reports on the trending ‘echo’. \r\nDescription taken from paper: \r\nArviv, E., Hanouna, S., & Tsur, O. (2020). It's a Thin Line Between Love and Hate: Using the Echo in Modeling Dynamics of Racist Online Communities. ArXiv, abs/2012.01133.","description_withheld":null,"homepage":"https://www.naslab.ise.bgu.ac.il/","introduced_date":"2020-11-16","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Hate Speech Detection","url":"/task/hate-speech-detection","datasets_with_task":"/datasets/task/hate-speech-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Echo Corpus"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}