Datasets › UMass Citation Field Extraction

UMass Citation Field Extraction

10 May 2013 archive 2025-07-28

The University of Massachusetts Amherst citation field extraction dataset contains labels and segments for extracted citations from articles found on arXiv. Compared to previous standard datasets in citation field extraction, this one had 4 times more data and provided detailed nested labels rather than coarse-grained flat labels, alongside drawing from 4 different academic disciplines versus 1 - namely computer science, mathematics, physics, and quantitative biology.

It consisted of 6,000 unlabeled citation strings, with 1829 labeled to date at the time of its last publication - 2476 according to the latest citation from 'Using BibTeX to Automatically Generate Labeled Data for Citation Field Extraction' - Dung Thai, Zhiyang Xu, Nicholas Monath, Boris Veytsman, and Andrew McCallum. Each citation string was labeled hierarchically, separating coarse-grain and fine-grain labeled segments.

Dataset introduced in the following paper:

Sam Anzaroot and Andrew McCallum. A new dataset for fine-grained citation field extraction. In ICML Workshop on Peer Reviewing and Publishing Models (PEER), 2013.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

Unknown - TBC

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • UMass Citation Field Extraction

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections