Browse State-of-the-Art › Authorship Attribution
Authorship Attribution
62 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Authorship attribution, also known as authorship identification, aims to attribute a previously unseen text of unknown authorship to one of a set of known authors.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 62 papers with code (212 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Sep 2021 3 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecent progress in generative language models has enabled machines to generate astonishingly realistic texts.
-
21 Sep 2016 3 repositories listedConvolutional neural networks (CNNs) have demonstrated superior capability for extracting information from raw signals in computer vision.
-
24 Aug 2023 2 repositories listedIn this work, we use a multilingual knowledge distillation approach to train BERT models to produce sentence embeddings for Ancient Greek text.
-
30 Oct 2019 2 repositories listedWe apply this measure to authorship attribution challenges, where the goal is to identify the author of a document using other documents whose authorship is known.
-
18 Feb 2025 1 repository listedIn this paper, we propose to adjust the masking ratio and to decide which tokens to mask based on a novel task-informed anti-curriculum learning scheme.
-
22 Jan 2025 1 repository listedNatural Language Processing (NLP) for lesser-resourced languages faces persistent challenges, including limited datasets, inherited biases from high-resource languages, and the need for domain-specific solutions.
-
18 Dec 2024 1 repository listedHuman trafficking (HT) remains a critical issue, with traffickers increasingly leveraging online escort advertisements (ads) to advertise victims anonymously.
-
6 Dec 2024 1 repository listedIn an era where cyberattacks increasingly target the software supply chain, the ability to accurately attribute code authorship in binary files is critical to improving cybersecurity measures.
-
6 Oct 2024 1 repository listedThis study introduces a novel dataset derived from the Medical Informatics Europe (MIE) Conference proceedings, addressing the need for sophisticated analytical tools in the field.
-
12 Aug 2024 1 repository listedHowever, current AA benchmarks commonly overlook this uniqueness and frame the problem as a closed-world classification, assuming a fixed number of authors throughout the system's lifespan and neglecting the inclusion…
-
12 Jun 2024 1 repository listedAccordingly, we propose a Multi-task Figurative Language Model (MFLM) that learns to detect multiple FL features in text at once.
-
28 May 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)However, this information is underutilized as the weights are uninterpretable, and publicly available models are disorganized.
-
2 May 2024 1 repository listedSuch an approach, however, is inadequate for the whistleblowing scenario since it neglects further re-identification potential in textual features, including writing style.
-
13 Mar 2024 1 repository listed(3) Can LLMs provide explainability in authorship analysis, particularly through the role of linguistic features?
-
1 Feb 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)AO is the corresponding adversarial task, aiming to modify a text in such a way that its semantics are preserved, yet an AA model cannot correctly infer its authorship.
-
2 Dec 2023 1 repository listedIn this technical note we suggest a novel approach to discover temporal (related and unrelated to language dilation) and personality (authorship attribution) aspects in historical datasets.
-
13 Nov 2023 1 repository listedAuthorship verification is the task of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text.
-
3 Nov 2023 1 repository listedWhile a substantial amount of work has recently been devoted to enhance the performance of computational Authorship Identification (AId) systems, little to no attention has been paid to endowing AId systems with the…
-
25 Oct 2023 1 repository listedThe large language based-model chatbot ChatGPT gained a lot of popularity since its launch and has been used in a wide range of situations.
-
9 Oct 2023 1 repository listedWe examine if this UID principle can help capture differences between Large Language Models (LLMs)-generated and human-generated texts.
-
30 Sep 2023 1 repository listedMoreover, 'People of Yue, Forget Not Your Ancestors' Instructions' seems to be either predominantly authored or extensively revised by Lu Xun given its notable stylistic similarities to 'Looking at the Land of Yue,' a…
-
22 Sep 2023 1 repository listedWe propose TopFormer to improve existing AA solutions by capturing more linguistic patterns in deepfake texts by including a Topological Data Analysis (TDA) layer in the Transformer-based model.
-
15 Sep 2023 1 repository listedVia automatic and human evaluation, we show that specialized student models fine-tuned on our datasets outperform generalist teacher models on the explainable style transfer task in one-shot settings, and perform…
-
22 Aug 2023 1 repository listedAutomatically disentangling an author's style from the content of their writing is a longstanding and possibly insurmountable problem in computational linguistics.
-
14 Aug 2023 1 repository listedLarge language models (LLMs) such as GPT-4, PaLM, and Llama have significantly propelled the generation of AI-crafted text.
-
1 Apr 2023 1 repository listedAutomatic Authorship Attribution (AAA) is the result of applying tools and techniques from Digital Humanities to authorship attribution studies.
-
24 Jan 2023 1 repository listedWe investigate the effects on authorship identification tasks of a fundamental shift in how to conceive the vectorial representations of documents that are given as input to a supervised learner.
-
19 Dec 2022 1 repository listedIn addition to being a widely recognised novelist, Milan Kundera has also authored three pieces for theatre: The Owners of the Keys (Majitel\'e kl\'i\v{c}\r{u}, 1961), The Blunder (Pt\'akovina, 1967), and Jacques and…
-
14 Nov 2022 1 repository listedIn this work, we present a transformer-based, neural-network architecture that only uses the text content and the author names in the bibliography to attribute an anonymous manuscript to an author.
-
9 Nov 2022 1 repository listedDetermining the author of a text is a difficult task.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections