Browse State-of-the-Art › Document Classification
Document Classification
235 papers with code · 21 benchmarks · 18 datasets archive 2025-07-28
Document Classification is a procedure of assigning one or more labels to a document from a predetermined set of labels.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
22 leaderboard tables shown for this task, 21 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 22 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 235 papers with code (641 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Oct 2017 93 repositories listed Syntology ran 50 of 106 samples · 56 unverified · 43 pointer-only (licence)We present graph attention networks (GATs), novel neural network architectures that operate on graph-structured data, leveraging masked self-attentional layers to address the shortcomings of prior methods based on graph…
-
9 Sep 2016 55 repositories listed Syntology ran 31 of 58 samples · 27 unverified · 22 pointer-only (licence)We present a scalable approach for semi-supervised learning on graph-structured data that is based on an efficient variant of convolutional neural networks which operate directly on graphs.
-
29 Mar 2016 26 repositories listed Syntology ran 15 of 28 samples · 13 unverified · 4 pointer-only (licence)We present a semi-supervised learning framework based on graph embeddings.
-
14 Jun 2017 15 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 1 pointer-only (licence)Confidence calibration -- the problem of predicting probability estimates representative of the true correctness likelihood -- is important for classification models in many applications.
-
27 May 2022 13 repositories listed Syntology ran 9 of 30 samples · 21 unverified · 1 pointer-only (licence)We also extend FlashAttention to block-sparse attention, yielding an approximate attention algorithm that is faster than any existing approximate attention method.
-
26 Dec 2018 13 repositories listed Syntology ran 4 of 10 samples · 6 unverified · 4 pointer-only (licence)We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts.
-
11 Jun 2018 13 repositories listedWe demonstrate that large gains on these tasks can be realized by generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task.
-
2 Nov 2019 7 repositories listed Syntology ran 7 of 35 samples · 28 unverifiedMoreover, it is shown that reasonable performance can be obtained when ZEN is trained on a small corpus, which is important for applying pre-training techniques to scenarios with limited data.
-
15 Apr 2020 5 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph.
-
19 Oct 2022 4 repositories listedPre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain.
-
22 Mar 2020 4 repositories listedIn this paper, we model the problem of finding the relationship between two documents as a pairwise document classification task.
-
10 Sep 2019 4 repositories listedPretrained language models are promising particularly for low-resource languages as they only require unlabelled data.
-
13 Jun 2019 4 repositories listed Syntology ran 0 of 2 samples · 2 unverified
-
23 Apr 2017 4 repositories listedRecurrent Neural Networks are showing much promise in many sub-areas of natural language processing, ranging from document classification to machine translation to automatic question answering.
-
25 Nov 2016 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedRecently, there has been an increasing interest in geometric deep learning, attempting to generalize deep learning methods to non-Euclidean structured data such as graphs and manifolds, with a variety of applications…
-
1 Oct 2014 4 repositories listed
-
17 Jun 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 1 pointer-only (licence)Scientific documents record research findings and valuable human knowledge, comprising a vast corpus of high-quality data.
-
9 Dec 2019 3 repositories listedPage classification is a crucial component to any document analysis system, allowing for complex branching control flows for different components of a given document.
-
23 Oct 2019 3 repositories listedBERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm.
-
15 Jul 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Classification of document images is a critical step for archival of old manuscripts, online subscription and administrative procedures.
-
17 Apr 2019 3 repositories listed Syntology ran 0 of 14 samples · 14 unverifiedWe present, to our knowledge, the first application of BERT to document classification.
-
2 Nov 2018 3 repositories listedSpecifically, we use bi-directional Recurrent Neural Networks, together with max-pooling over the temporal/sequential dimension and neural attention, for representing (i) the headline, (ii) the first two sentences of…
-
24 Sep 2017 3 repositories listedThis is because along with this growth in the number of documents has come an increase in the number of categories.
-
6 Feb 2024 2 repositories listedHowever, evaluating GLLMs presents a challenge as the binary true or false evaluation used for discriminative models is not applicable to the predictions made by GLLMs.
-
10 Sep 2021 2 repositories listedHere, we introduce the application of balancing loss functions for multi-label text classification.
-
2 Aug 2021 2 repositories listedEuroVoc is a multilingual thesaurus that was built for organizing the legislative documentary of the European Union institutions.
-
16 Apr 2021 2 repositories listedToken-level analysis shows that temporal adaptation captures event-driven changes in language use in the downstream task, but not those changes that are actually relevant to task performance.
-
18 Jan 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn this work we study a mathematical formalization of this network motif and apply it to learning the correlational structure between words and their context in a corpus of unstructured text, a common natural language…
-
14 Jan 2021 2 repositories listedOur approach substantially outperforms previous results on top-50 medical code prediction on MIMIC-III dataset.
-
14 Oct 2020 2 repositories listedIn this paper, we explore the potential of only using the label name of each class to train classification models on unlabeled data, without using any labeled documents.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections