Papers › A Tidy Data Model for Natural Language Processing using cleanNLP

A Tidy Data Model for Natural Language Processing using cleanNLP

27 Mar 2017arXiv:1703.09570archive 2025-07-28

Taylor Arnold

The package cleanNLP provides a set of fast tools for converting a textual corpus into a set of normalized tables. The underlying natural language processing pipeline utilizes Stanford's CoreNLP library, exposing a number of annotation tasks for text written in English, French, German, and Spanish. Annotators include tokenization, part of speech tagging, named entity recognition, entity linking, sentiment analysis, dependency parsing, coreference resolution, and information extraction.

PaperPDFCode

Code

statsmaths/cleanNLP officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Coreference ResolutionDependency ParsingEntity LinkingNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech TaggingSentiment Analysiscoreference-resolutionnamed-entity-recognition

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections