Papers › EDGAR-CORPUS: Billions of Tokens Make The World Go Round

EDGAR-CORPUS: Billions of Tokens Make The World Go Round

29 Sep 2021EMNLP (ECONLP) 2021 11arXiv:2109.14394archive 2025-07-28

Lefteris Loukas, Manos Fergadiotis, Ion Androutsopoulos, Prodromos Malakasiotis

We release EDGAR-CORPUS, a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years. To the best of our knowledge, EDGAR-CORPUS is the largest financial NLP corpus available to date. All the reports are downloaded, split into their corresponding items (sections), and provided in a clean, easy-to-use JSON format. We use EDGAR-CORPUS to train and release EDGAR-W2V, which are WORD2VEC embeddings for the financial domain. We employ these embeddings in a battery of financial NLP tasks and showcase their superiority over generic GloVe embeddings and other existing financial word embeddings. We also open-source EDGAR-CRAWLER, a toolkit that facilitates downloading and extracting future annual reports.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Word Embeddings

Datasets

Introduced by this paper, per the archive.

EDGAR-CORPUS

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

GloVe

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections