{"url":"/dataset/trecdd","name":"TRECDD","full_name":"TREC Dynamic Domain","description_markdown":"The dataset used for TREC 2017 Dynamic Domain Track consists of two domains: Ebola and New York Times.\r\n\r\n1.1 Ebola\r\n\r\nThe Ebola dataset is crawled by Juliana Friere (NYU, juliana dot freire at nyu dot edu), Kien Pham(NYU), Peter Landwehr (Giant Oak, peter dot landwehr at giantoak dot com) and Lewis McGibbney (JPL, Lewis dot J dot Mcgibbney at jpl dor nasa dot gov).\r\n\r\nThe Ebola dataset contains records related to the Ebola outbreak in Africa in 2014-2015. The original dataset includes tweets relating to the outbreak, web pages from sites hosted in the affected countries as well as PDF documents from websites such as World Health Organization, Financial Tracking Service and The World Bank. Such information resources are designed to provide information to citizens and aid workers on the ground.\r\n\r\n1.2 New York Times\r\n\r\nThe New York Times dataset is published by Evan Sandhaus in 2008 under LDC Catalog No. LDC2008T19.\r\n\r\nThe New York Times dataset consists of articles published in New York Times from January 1, 1987 to June 19, 2007 with metadata provided by the New York Times Newsroom, the New York Times Indexing Service and the online production staff at nytimes.com. Most articles are manually summarized and tagged by professional staffs. The original form of this dataset is in News Industry Text Format (NITF). This dataset can aid the research in Document Categorization, Information Retrieval, Entity Extraction and etc.","description_withheld":null,"homepage":"http://infosense.cs.georgetown.edu/trec_dd/dataset.html","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"Custom","url":"http://infosense.cs.georgetown.edu/trec_dd/dataset.html"},"modalities":[],"tasks":[],"languages":[],"variants":["TRECDD"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}