Papers › GROBID: Combining Automatic Bibliographic Data Recognition and Term Extraction for...

GROBID: Combining Automatic Bibliographic Data Recognition and Term Extraction for Scholarship Publications

1 Sep 2009Research and Advanced Technology for Digital Libraries 2009 9archive 2025-07-28

Patrice Lopez

Based on state of the art machine learning techniques, GROBID (GeneRation Of BIbliographic Data) performs reliable bibliographic data extractions from scholar articles combined with multi-level term extractions. These two types of extraction present synergies and correspond to complementary descriptions of an article. This tool is viewed as a component for enhancing the existing and the future large repositories of technical and scientific publications.

PaperPDFCode

Code

kermitt2/grobid Apache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ArticlesTerm Extraction

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections