Datasets › Labeled data for citation field extraction
Labeled data for citation field extraction
Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science. Citation field extraction is the task of parsing citations: given a citation string, extract authors, title, venue, doi etc. Since the number of citations is counted by hundreds millions, efficient computer based methods for this task are very important.
The development of machine learning methods for citation field extraction requires ground truth: a large corpus of labeled citations. This dataset provides a very large (41M) corpus of labeled data obtained by the reverse process: we took structured citation lists and used BibTeX to generate labeled citation strings.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
CC0
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Labeled data for citation field extraction
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections