Browse State-of-the-Art › Binary text classification
Binary text classification
12 papers with code · 7 benchmarks · 10 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Ghostbuster (All Domains) (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| MAGE (Arbitrary-domains & Arbitrary-models) (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| MixSet (Binary) (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| TURINGBENCH (Turing Test, GPT-3) (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| TURINGBENCH (Turing Test, FAIR_wmt20) (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| TweepFake (2 rows) | GigaCheck (Mistral-7B) | GigaCheck: Detecting LLM-generated Content | — | — | Compare |
| ECHR Non-Anonymized (1 row) | HIER-BERT | Neural Legal Judgment Prediction in English | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
12 shown of 12 papers with code (20 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Sep 2021 3 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecent progress in generative language models has enabled machines to generate astonishingly realistic texts.
-
11 Jan 2024 2 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 16 pointer-only (licence)With the rapid development and widespread application of Large Language Models (LLMs), the use of Machine-Generated Text (MGT) has become increasingly common, bringing with it potential risks, especially in terms of…
-
8 Jun 2023 2 repositories listedIn this article, we present DACCORD, a new dataset dedicated to the task of automatically detecting contradictions between sentences in French.
-
24 May 2023 2 repositories listedIn conjunction with our model, we release three new datasets of human- and AI-generated text as detection benchmarks in the domains of student essays, creative writing, and news articles.
-
22 May 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedIn practical scenarios, however, the detector faces texts from various domains or LLMs without knowing their sources.
-
20 Jul 2018 2 repositories listedThe disambiguation is formulated as a binary text classification problem where the prediction is made for the potential soft skill based on the context where it occurs.
-
1 Dec 2014 2 repositories listedUsing online EP and the central limit theorem we find an analytical approximation to the Bayes update of this posterior, as well as the resulting Bayes estimates of the weights and outputs.
-
3 May 2024 1 repository listedHoaxes are a recognised form of disinformation created deliberately, with potential serious implications in the credibility of reference knowledge resources such as Wikipedia.
-
Identification of the Relevance of Comments in Codes Using Bag of Words and Transformer Based Models11 Aug 2023 1 repository listedThe performance of the classical bag of words model and transformer-based models were explored to identify significant features from the given training corpus.
-
1 Nov 2020 1 repository listedHere, we present a large-scale empirical study on active learning techniques for BERT-based classification, addressing a diverse set of AL strategies and datasets.
-
31 Jul 2020 1 repository listedTo prevent this, it is crucial to develop deepfake social media messages detection systems.
-
12 Sep 2019 1 repository listedShallow machine learning strategies showed lower overall micro F1 scores, but still higher than deep learning strategies and the baseline.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections