Home › Code › clean_text

clean_text

Syntologyentry name in harvested coderead from the graph 2026-09-24

clean_text appears in the code Syntology harvested for 84 papers, as 80 distinct code bodies found in 87 places (a place is one code body under one paper). At least one of them ran in 34 of the papers; 29 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named clean_text do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 30 of the 80 distinct code bodies named clean_text; 50 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
11ran · our draft was wrong
0ran · fixture could not drive it
19ran
50unverified
29fingerprinted

Licence is a property of each copy, so it is counted per place: 29 of the 87 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

84 papers shown of 84, newest first; 87 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 21 papers added by Syntology; 6 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
On the Design Fundamentals of Pixel Text Representation Learning added by Syntology 2026-09 (from id) Pixel-Linguist/Pixel-Linguist-II/training/filter_dataset.py b5bb3eecb46f7556 unverified Apache-2.0 (permissive)
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes added by Syntology 2026-08 (from id) heathriel/synthetic-memoir-audit/analysis/09_build_public_source_registry.py 291a30968bfd353d ran fingerprinted no licence file found · pointer only
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options added by Syntology 2026-08 (from id) obedjunias19/structured-compositional-reasoning/lsata/run_structured_inference.py b26dd53b8f714434 ran fingerprinted MIT (permissive)
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs added by Syntology 2026-07 (from id) ECLADATTA/KONTRAST/Output_GRASP/Qwen3-4B-Instruct-2507/ComplexQA/tag.py cc20d3e7db14a81f ran · our draft was wrong fingerprinted no licence file found · pointer only
The Cross-Domain Generalization Cost of Offensive Language Detection added by Syntology 2026-07 (from id) renruixing/The-Cross-Domain-Generalization-Cost-of-Offensive-Language-Detection/prepare_dataset.py 2b3dcbbecbaa9c5b ran fingerprinted no licence file found · pointer only
Automated Discovery Has No Universally Superior Harness added by Syntology 2026-07 (from id) akshat57/harness-generalization/allocation/plot_sheet_config_score_boxes.py 28dcdd8e4cc20e02 ran fingerprinted MIT (permissive)
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling added by Syntology 2026-06 (from id) kaist-cvml/perception-judge/prepare-datasets/post_processing_filter_after_gen.py cfeb19d54f2339b3 ran fingerprinted no licence file found · pointer only
ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering added by Syntology 2026-05 (from id) yikai-zhu/ConRAG/src/conrag/common.py e0d187ce27d64511 unverified MIT (permissive)
Sentiment Analysis of Mobile Legends App Reviews Using Machine Learning and LSTM-Based Deep Learning Models added by Syntology 2026-05 (from id) Viramhrani/pba2026-Kelompok16/app/app_dl.py 06580d402622ea57 ran fingerprinted no licence file found · pointer only
Sentiment and Emotion Classification of Indonesian E-Commerce Reviews via Multi-Task BiLSTM and AutoML Benchmarking added by Syntology 2026-04 (from id) ikii-sd/pba2026-crazyrichteam/src/preprocessing.py a0244a869f42339c unverified no licence file found · pointer only
Enhancing Unsupervised Keyword Extraction in Academic Papers through Integrating Highlights with Abstract added by Syntology 2026-04 (from id) xiangyi-njust/Highlight-KPE/code/MDERank/mderank.py 276c875f8460a41e unverified no licence file found · pointer only
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection added by Syntology 2026-04 (from id) thinkst/zippy/zippy/zippy.py 279e23ca5da96480 unverified MIT (permissive)
AOP-Smart: A RAG-Enhanced Large Language Model Framework for Adverse Outcome Pathway Analysis added by Syntology 2026-04 (from id) qinjiang-lab/AOP-Smart/AOP-Smart.py b08f75a14f197c80 unverified MIT (permissive)
AOP-Smart: A RAG-Enhanced Large Language Model Framework for Adverse Outcome Pathway Analysis added by Syntology 2026-04 (from id) qinjiang-lab/AOP-Smart/XML_analysis.py 0ddb4aa5850a0592 unverified MIT (permissive)
Drift and selection in LLM text ecosystems added by Syntology 2026-04 (from id) SR123/LLM-text-ecosystems/src/drift_selection/cleaning.py 1b6456251fc3b160 unverified MIT (permissive)
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection added by Syntology 2026-03 (from id) Nouran-Khallaf/denoising-difficulty/models/Baseline.py e9eb5d0c4765858b unverified no licence file found · pointer only
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection added by Syntology 2026-03 (from id) Nouran-Khallaf/denoising-difficulty/models/baseline.py 332ce80bbbda4118 unverified no licence file found · pointer only
Differences in Typological Alignment in Language Models' Treatment of Differential Argument Marking added by Syntology 2026-02 (from id) Iskar-Deng/DAM-learning/data_processing/parse.py 42072772adaf6878 unverified no licence file found · pointer only
MultiCube-RAG for Multi-hop Question Answering added by Syntology 2026-02 (from id) OpenBMB/UltraRAG/servers/corpus/src/corpus.py 4de9e8ee4cbe4329 unverified Apache-2.0 (permissive)
Q-Hawkeye: Reliable Visual Policy Optimization for Image Quality Assessment added by Syntology 30 Jan 2026 AMAP-ML/Q-Hawkeye/src/virft/src/open_r1/grpo_am.py 7ed1cffd12a5fc04 unverified no licence file found · pointer only
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning added by Syntology 2026-01 (from id) Multimodal-Intelligence-Lab-MIL/MARS/code/MARS/gemini2.5flash.py e609dfa06255ac50 ran · our draft was wrong fingerprinted MIT (permissive)
Quantifying Climate Policy Action and Its Links to Development Outcomes: A Cross-National Data-Driven Analysis added by Syntology 2025-10 (from id) booktrackerGirl/climate_change_policy_analysis/src/extract_content.py 818185916e2d10a0 unverified MIT (permissive)
CORRECT: Condensed Error Recognition via Knowledge Transfer in Multi-agent Systems added by Syntology 2025-09 (from id) UIUC-MLSys/CORRECT/src/Lib/cloud_paper.py 18f2cdd35c9ac991 unverified licence not identified · pointer only
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions 8 Jul 2025 lukahhcm/cultureclip/data_curation/diverse_caption.py 98f3d9a2dbe966f8 unverified no licence file found · pointer only
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations 20 May 2025 persona-bench/persona/PromptMakers/3.3PromptMaker.py b3a8c68af7fbac6f ran · our draft was wrong fingerprinted no licence file found · pointer only
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings 19 Mar 2025 ny1024/deepseek-safety-eval/eval/base/deepseek_r1.py 618e93855f6e4e2b ran · our draft was wrong fingerprinted no licence file found · pointer only
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond 13 Mar 2025 Qihoo360/Light-R1/decontaminate/character_matching.py 61619ba901c7b574 unverified Apache-2.0 (permissive)
The Rise of AI-Generated Content in Wikipedia 10 Oct 2024 brooksca3/wiki_collection/eval/binoculars/run_wiki_binoculars.py 722efc24678ccc7e ran fingerprinted no licence file found · pointer only
Style-Specific Neurons for Steering LLMs in Text Style Transfer 1 Oct 2024 wenlai-lavine/sNeuron-TST/Evaluation/cls/authorship.py 5eb4d0e1f153b895 ran fingerprinted no licence file found · pointer only
Cottention: Linear Transformers With Cosine Attention 27 Sep 2024 gmongaras/Cottention_Transformer/BERT_Trainer/create_hf_datasets.py 774558f531286112 unverified no licence file found · pointer only
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage 17 Sep 2024 osu-nlp-group/eia_against_webagent/SeeAct/src/data_utils/dom_utils.py 0e07e2733eab1746 ran fingerprinted MIT (permissive)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding 28 Aug 2024 dptech-corp/Uni-SMART/SciLitLLM/sft/helper/parse_pdfs.py 2e366df93e0e0280 ran fingerprinted MIT (permissive)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations 17 Jun 2024 Sri-Harsha/scpp/generate_sugarcrepe_plus-mistral.py 7c6b861ca90aa566 ran · our draft was wrong fingerprinted MIT (permissive)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation 7 Jun 2024 meta-inf/suq-nlg/qa/vu.py 2745a7b067527cea ran fingerprinted MIT (permissive)
Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space 22 May 2024 cognizant-ai-labs/semantic-density-paper/experiment_code/get_semantic_density_full_beam_search_unique_datasets_temperature.py 382e62efddba975b ran fingerprinted licence not identified · pointer only
Bottleneck-Minimal Indexing for Generative Document Retrieval 12 May 2024 kduxin/Bottleneck-Minimal-Indexing/NCIRetriever/model.py 5d140ce18f0b48ed ran fingerprinted MIT (permissive)
Keep It Private: Unsupervised Privatization of Online Text 16 May 2024 csbao/kip-privatization/src/generator.py 32acbc36604380fe ran fingerprinted no licence file found · pointer only
RULER: What's the Real Context Size of Your Long-Context Language Models? 9 Apr 2024 mungg/OneRuler/OneRuler/eval/evaluate.py d7f2c747adfac407 ran · our draft was wrong fingerprinted MIT (permissive)
Learning to Plan for Language Modeling from Unlabeled Data 31 Mar 2024 natithan/learning-to-plan-for-language-modeling-from-unlabeled-data/eval_generations.py efc543ba6ad34128 ran fingerprinted no licence file found · pointer only
Language Repository for Long Video Understanding 21 Mar 2024 kkahatapitiya/langrepo/util.py d3cd8c8e5b09319c ran fingerprinted MIT (permissive)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models 11 Mar 2024 pkunlp-icler/fastv/src/FastV/inference/eval/inference_aokvqa.py 22967b94d546f593 unverified no licence file found · pointer only
On the Multi-turn Instruction Following for Conversational Web Agents 23 Feb 2024 magicgh/self-map/src/dom_utils.py 0e07e2733eab1746 ran fingerprinted MIT (permissive)
GPT-4V(ision) is a Generalist Web Agent, if Grounded 3 Jan 2024 osu-nlp-group/seeact/src/data_utils/dom_utils.py 0e07e2733eab1746 ran fingerprinted no licence file found · pointer only
Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement 27 Oct 2023 rosthalken/legal-interpretation/create_samples.py b305b281072ece66 ran fingerprinted no licence file found · pointer only
TiC-CLIP: Continual Training of CLIP Models 24 Oct 2023 apple/ml-tic-clip/dataset_creation/tic-yfcc15m/create_splits.py 9b270ee2a7164544 ran fingerprinted licence not identified · pointer only
Contextual Label Projection for Cross-Lingual Structured Prediction 16 Sep 2023 pluslabnlp/clap/src/utils.py ef707ceb5ff307d7 unverified no licence file found · pointer only
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context 15 Sep 2023 k2-fsa/libriheavy/scripts/extract_and_normalize_transcript.py 20b4ccb36cf358ec ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing 2 Sep 2023 cwang621/blsp/data_process/prepare_alpaca.py 61134cad531c377a unverified Apache-2.0 (permissive)
On the Efficacy of Sampling Adapters 7 Jul 2023 rycolab/sampling-adapters/utils.py 25466aac44e0d840 ran no licence file found · pointer only
Recurrent Attention Networks for Long-text Modeling 12 Jun 2023 4ai/ran/examples/LongCLF/run_20ng.py 08092de6a633d0d8 ran · our draft was wrong fingerprinted MIT (permissive)
Task-Optimized Adapters for an End-to-End Task-Oriented Dialogue System 4 May 2023 sogang-isds/TOATOD/E2E_TOD/clean_dataset.py 066f7f2b00042eec unverified Apache-2.0 (permissive)
No more Reviewer #2: Subverting Automatic Paper-Reviewer Assignment using Adversarial Learning 25 Mar 2023 rub-syssec/adversarial-papers/src/utils/attack_utils.py 0f71c17baa894982 unverified MIT (permissive)
Guiding Large Language Models via Directional Stimulus Prompting 22 Feb 2023 leezekun/directional-stimulus-prompting/sft4lms/MultiWOZ/clean_dataset.py ea3e5b726884db6d unverified Apache-2.0 (permissive)
Benchmarking Large Language Models for Automated Verilog RTL Code Generation 13 Dec 2022 shailja-thakur/vgen/pdf_extraction_instance.py 66956cac853d371d unverified Apache-2.0 (permissive)
A Generative User Simulator with GPT-based Architecture and Goal State Tracking for Reinforced Multi-Domain Dialog Systems 17 Oct 2022 thu-spmi/gus/clean_dataset.py c20b24a9790d8f15 unverified Apache-2.0 (permissive)
Re2G: Retrieve, Rerank, Generate 13 Jul 2022 ibm/kgi-slot-filling/slot_filling/kilt_passage_corpus.py d890771bd3e4b3f7 unverified Apache-2.0 (permissive)
Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners 22 May 2022 mikewangwzhl/vidil/eval_video_captioning_results.py 1ff44f0d81c9de63 unverified MIT (permissive)
Self-attention Does Not Need $O(n^2)$ Memory 10 Dec 2021 X-iZhang/Libra/libra/eval/radiology_report.py 46c06fec5aedf4a4 unverified Apache-2.0 (permissive)
Self-attention Does Not Need $O(n^2)$ Memory 10 Dec 2021 X-iZhang/Libra/libra/eval/temporal_f1.py cb00b798f297e5f3 unverified Apache-2.0 (permissive)
Unifying Multimodal Transformer for Bi-directional Image and Text Generation 19 Oct 2021 researchmm/generate-it/diverse-it-generator/sample_images.py 54027d44ffd92b3d unverified MIT (permissive)
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System 29 Sep 2021 awslabs/pptod/E2E_TOD/clean_dataset.py ea3e5b726884db6d unverified Apache-2.0 (permissive)
Robust Retrieval Augmented Generation for Zero-shot Slot Filling 31 Aug 2021 IBM/retrieve-write-slot-filling/slot_filling/kilt_passage_corpus.py d890771bd3e4b3f7 unverified Apache-2.0 (permissive)
Summary Explorer: Visualizing the State of the Art in Text Summarization 4 Aug 2021 webis-de/summary-explorer/text-processing/src/utils.py f7918be2e65b31a7 unverified MIT (permissive)
$Q^{2}$: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering 16 Apr 2021 orhonovich/q-squared/pipeline/score.py 96af2fcff56e897c unverified Apache-2.0 (permissive)
SummVis: Interactive Visual Analysis of Models, Data, and Evaluation for Text Summarization 15 Apr 2021 robustness-gym/summvis/utils.py 7c9b5b54376420a1 unverified Apache-2.0 (permissive)
MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue Systems 25 Sep 2020 zlinao/MinTL/damd_multiwoz/clean_dataset.py 5b9d615f09edd9c8 unverified MIT (permissive)
DocVQA: A Dataset for VQA on Document Images 1 Jul 2020 anisha2102/docvqa/create_dataset.py 58efca4ff69d4f1c unverified MIT (permissive)
Improving GAN Training with Probability Ratio Clipping and Sample Reweighting 12 Jun 2020 Holmeswww/PPOGAN/style_transfer/prepare_manual.py d8ab4cdbe05b528e unverified MIT (permissive)
Libri-Light: A Benchmark for ASR with Limited or No Supervision 17 Dec 2019 identical code first harvested elsewhere 20b4ccb36cf358ec ran · our draft was wrong fingerprinted licence of this copy not recorded
Using Clinical Notes with Time Series Data for ICU Management 12 Sep 2019 kaggarwal/ClinicalNotesICU/scripts/extract_notes.py 9c3d3e9f83b541b3 ran · our draft was wrong fingerprinted MIT (permissive)
Multilingual and Multi-Aspect Hate Speech Analysis 29 Aug 2019 HKUST-KnowComp/MLMA_hate_speech/annotated_data_processing.py bfe8c53bc293a85f unverified MIT (permissive)
Auditing Radicalization Pathways on YouTube 2019-08 (from id) markledwich2/YouTubeNetworks/DataScripts/video_entities.py 37c85cb881dd4a3e unverified MIT (permissive)
Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites 19 Aug 2019 VAShibaev/text_style_transfer/shiftedae/prepare_manual.py d8ab4cdbe05b528e unverified Apache-2.0 (permissive)
ReQA: An Evaluation for End-to-End Answer Retrieval Models 10 Jul 2019 google/retrieval-qa-eval/nq_to_squad.py 0b0ded1135f95b47 unverified Apache-2.0 (permissive)
NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language 14 Jun 2019 leonweber/nlprolog/preprocessing.py 5dc0130fdfd31748 unverified MIT (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 jsantoso2/yelp-clone-ml-project/app/flask-backend/firebase.py c7f933f28568cfaa unverified Apache-2.0 (permissive)
Deep Residual Learning for Small-Footprint Keyword Spotting 28 Oct 2017 magahub/honk/utils/client.py 259debd1d954e429 unverified MIT (permissive)
Neural Models for Documents with Metadata 25 May 2017 dallascard/scholar/preprocess_data.py 128fcff9d8fad853 ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
FastText.zip: Compressing text classification models 12 Dec 2016 currentsapi/fastlangid/fastlangid/utils.py cc21a5b05edde91f unverified Apache-2.0 (permissive)
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin 8 Dec 2015 raraz15/DeepTurkish/utilities/text_format.py 1486f8a5a335a589 unverified MIT (permissive)
Sequence to Sequence Learning with Neural Networks 10 Sep 2014 yash-nishaant/Seq2Seq-Chatbot/chatbot.py 8eaa3f3c2fa1ed5d ran · our draft was wrong fingerprinted no licence file found · pointer only
arXiv:ijcai2024_0687 zjunlp/FactCHD/data_generate/text_data_generate.py 4f2d8c4335e3a017 unverified MIT (permissive)
arXiv:2025.findings-acl.294 kkahatapitiya/LangRepo/util.py d3cd8c8e5b09319c ran fingerprinted MIT (permissive)
arXiv:2023.findings-emnlp.730 chufeiluo/legalhatespeech/zeroshot/generate_hf.py 03cdc84617a84fa0 unverified MIT (permissive)
arXiv:2023.findings-emnlp.384 JHL-HUST/SparseMA/dataloader/crisismmd_dataset.py b8be2ec62a547e2a unverified Apache-2.0 (permissive)
arXiv:2023.findings-acl.416 cxa-unique/IDEM/collect_surrogate_train_data_nq.py 84895c4785f47e06 unverified Apache-2.0 (permissive)
arXiv:2021.findings-emnlp.112 bepoetree/MTTOD/utils/clean_dataset.py d7e520c85b455eb3 unverified Apache-2.0 (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections