{"url":"/dataset/neuralnews","name":"NeuralNews","full_name":null,"description_markdown":"NeuralNews is a dataset for machine-generated news detection. It consists of human-generated and machine-generated articles. The human-generated articles are extracted from the GoodNews dataset, which is extracted from the New York Times. It contains 4 types of articles:\r\n\r\n- Real Articles and Real Captions\r\n- Real Articles and Generated Captions\r\n- Generated Articles and Real Captions\r\n- Generated Articles and Generated Captions\r\n\r\nIn total, it contains about 32K samples of each article type (resulting in about 128K total).\r\n\r\nSource: [Detecting Cross-Modal Inconsistency to Defend Against Neural Fake News](https://arxiv.org/abs/2009.07698)","description_withheld":null,"homepage":"https://cs-people.bu.edu/rxtan/projects/didan/","introduced_date":"2020-09-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/detecting-cross-modal-inconsistency-to-defend","title":"Detecting Cross-Modal Inconsistency to Defend Against Neural Fake News","first_author":"Reuben Tan","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NeuralNews"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}