Datasets › New York Times Annotated Corpus

New York Times Annotated Corpus

Introduced in The New York Times Annotated Corpus17 Oct 2008 archive 2025-07-28

The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times Indexing Service and the online production staff at nytimes.com. The corpus includes:

  • Over 1.8 million articles (excluding wire services articles that appeared during the covered period).
  • Over 650,000 article summaries written by library scientists.
  • Over 1,500,000 articles manually tagged by library scientists with tags drawn from a normalized indexing vocabulary of people, organizations, locations and topic descriptors.
  • Over 275,000 algorithmically-tagged articles that have been hand verified by the online production staff at nytimes.com. As part of the New York Times' indexing procedures, most articles are manually summarized and tagged by a staff of library scientists. This collection contains over 650,000 article-summary pairs which may prove to be useful in the development and evaluation of algorithms for automated document summarization. Also, over 1.5 million documents have at least one tag. Articles are tagged for persons, places, organizations, titles and topics using a controlled vocabulary that is applied consistently across articles. For instance if one article mentions "Bill Clinton" and another refers to "President William Jefferson Clinton", both articles will be tagged with "CLINTON, BILL".

Source: https://catalog.ldc.upenn.edu/LDC2008T19

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

25 shown of 25 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 262. The Syntology column is from Syntology's graph (read 2026-09-25), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Hierarchy-aware Biased Bound Margin Loss Function for Hierarchical Text Classification 1 4 13 Aug 2024 not harvested
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget 2 3 31 Jul 2024 not harvested
HILL: Hierarchy-aware Information Lossless Contrastive Learning for Hierarchical Text Classification 1 1 26 Mar 2024 ran 3 of 3 samples (0 unverified)
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction 1 1 12 Mar 2024 not harvested
HiTIN: Hierarchy-aware Tree Isomorphism Network for Hierarchical Text Classification 1 2 24 May 2023 ran 13 of 17 samples (4 unverified; 17 pointer-only for licence)
UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction 1 1 16 Nov 2022 not harvested
DeepStruct: Pretraining of Language Models for Structure Prediction 1 6 21 May 2022 ran 10 of 13 samples (3 unverified)
TDEER: An Efficient Translating Decoding Schema for Joint Extraction of Entities and Relations 1 2 1 Nov 2021 not harvested
REBEL: Relation Extraction By End-to-end Language generation 1 1 29 Oct 2021 not harvested
Zero-Shot Information Extraction as a Unified Text-to-Triple Translation 1 1 23 Sep 2021 not harvested
Adjacency List Oriented Relational Fact Extraction via Adaptive Multi-task Learning 1 1 3 Jun 2021 not harvested
KGPool: Dynamic Knowledge Graph Context Selection for Relation Extraction 1 1 1 Jun 2021 not harvested
Distantly-Supervised Long-Tailed Relation Extraction Using Constraint Graphs 1 1 24 May 2021 not harvested
Improving Distantly-Supervised Relation Extraction through BERT-based Label & Instance Embeddings 1 1 1 Feb 2021 not harvested
From Bag of Sentences to Document: Distantly Supervised Relation Extraction via Machine Reading Comprehension 1 1 8 Dec 2020 not harvested
Joint Entity and Relation Extraction with Set Prediction Networks 1 2 3 Nov 2020 not harvested
RECON: Relation Extraction using Knowledge Graph Context in a Graph Neural Network 1 1 18 Sep 2020 not harvested
Hierarchical Topic Mining via Joint Spherical Tree and Text Embedding 1 1 18 Jul 2020 not harvested
Dating Documents using Graph Convolution Networks 1 1 1 Feb 2019 not harvested
Neural Relation Extraction via Inner-Sentence Noise Reduction and Transfer Learning 0 2 21 Aug 2018 not harvested
Improving Distantly Supervised Relation Extraction using Word and Entity Based Attention 5 1 19 Apr 2018 not harvested
Neural Relation Extraction with Selective Attention over Instances 1 1 1 Aug 2016 not harvested
Distant Supervision for Relation Extraction via Piecewise Convolutional Neural Networks 1 1 1 Sep 2015 not harvested
A Burstiness-aware Approach for Document Dating 0 1 1 Jul 2014 not harvested
Labeling Documents with Timestamps: Learning from their Time Expressions 0 1 1 Jul 2012 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • New York Times Corpus
  • NYT
  • New York Times Annotated Corpus

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections