Home › Code › normalize_text

normalize_text

Syntologyentry name in harvested coderead from the graph 2026-09-24

normalize_text appears in the code Syntology harvested for 60 papers, as 54 distinct code bodies found in 61 places (a place is one code body under one paper). At least one of them ran in 27 of the papers; 24 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named normalize_text do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 24 of the 54 distinct code bodies named normalize_text; 30 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
7ran · our draft was wrong
0ran · fixture could not drive it
17ran
30unverified
24fingerprinted

Licence is a property of each copy, so it is counted per place: 18 of the 61 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

60 papers shown of 60, newest first; 61 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 24 papers added by Syntology; 11 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders added by Syntology 2026-09 (from id) atoz03/prepeft-probing/code/src/prepeft/eval/normalize.py 83355a93f4ab7fad unverified MIT (permissive)
DEBIASING TREE-BASED VARIABLE IMPORTANCE IN MIXED DATA added by Syntology 2026-09 (from id) melikechi-lab/just-add-noise/applications/mixed_importance_utils.py 8bbc006234ea32db unverified MIT (permissive)
Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation added by Syntology 2026-09 (from id) GIND123/CBCT-Clinical-Reasoner/src/cbct_reasoner/text.py fd60e97a9ffd66e8 unverified MIT (permissive)
T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation added by Syntology 2026-09 (from id) LLMSecResearch/T2LSC-Bench/evaluation/utils.py ee77ade1e729a87f unverified no licence file found · pointer only
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions added by Syntology 2026-08 (from id) QingfangLiu/llm-evidence-retrieval-bias/benchmark_tools/build_reference_indexing_from_cochrane_ris.py 20c568d253ca23b7 ran fingerprinted no licence file found · pointer only
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation added by Syntology 2026-08 (from id) NASK-NLP/PoVisLE/povisle/utils.py ecc81045ae9ebace ran fingerprinted Apache-2.0 (permissive)
M 3 R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding added by Syntology 2026-08 (from id) hongshi4/M3R-Bench/train/evaluate_sft_outputs.py 2b9ad881f4f0c4ac ran · our draft was wrong fingerprinted no licence file found · pointer only
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation added by Syntology 2026-08 (from id) ChenChiShui/FutureBridge-OPD/analysis/analyze_bridge_behavior.py f7d84e096670e78f unverified Apache-2.0 (permissive)
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean added by Syntology 2026-07 (from id) tlb-22/Euclean/src/codex_scheduler.py 55a9a37e10742ac4 ran fingerprinted Apache-2.0 (permissive)
Operational Proto-Introspection in Looped Language Models Process-Quality Taps, Executable Branching, and the Readout-Control Boundary added by Syntology 2026-07 (from id) VykosMolt/Branching-Looped-Transformer/probes/probe_layer_tap_cached_domains.py d8b04192866f751a ran fingerprinted Apache-2.0 (permissive)
Auditing Forgetting in Limited Memory Language Models added by Syntology 2026-07 (from id) raeesiarya/LMLMAudit/src/lmlm-audit/run_audit.py c1f14f57e9136ea6 ran · our draft was wrong fingerprinted MIT (permissive)
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? added by Syntology 2026-06 (from id) THU-KEG/RuVerBench/code/strategies/aggregate_voting_outputs.py ada78aa4aa08416f unverified licence not identified · pointer only
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent added by Syntology 2026-06 (from id) Morizeyao/EComAgentBench_/src/generation/utils.py 784e6a3cfa0b903a ran fingerprinted licence not identified · pointer only
DMF: A Deterministic Memory Framework for Conversational AI Agents added by Syntology 2026-06 (from id) matstech/dmf-benchmarks/dmf_bench/benchmarks/locomo/rigorous.py ae71bb9ba2af46bf ran fingerprinted MIT (permissive)
Automatic Layer Selection for Hallucination Detection added by Syntology 2026-05 (from id) DesoloYw/Automatic-Layer-Selection-for-Hallucination-Detection/src/metrics.py 4e686c57f3e71d5f ran fingerprinted no licence file found · pointer only
MasonNLP at MEDIQA-SYNUR 2026: Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction added by Syntology 2026-05 (from id) AHMRezaul/MEDIQA-SYNUR-2026/llama/rag.py cf405930dde4f507 ran · our draft was wrong fingerprinted no licence file found · pointer only
Knowledge Capsules: Structured Nonparametric Memory Units for LLMs added by Syntology 2026-04 (from id) jubin75/KVI/src/cleaning_and_dedupe.py 5aa9f2fbc850bd95 ran fingerprinted no licence file found · pointer only
Drift and selection in LLM text ecosystems added by Syntology 2026-04 (from id) SR123/LLM-text-ecosystems/src/drift_selection/cleaning.py 7987ba488637f64c unverified MIT (permissive)
Unified Deployment-Aware Evaluation of Open Reasoning Language Models added by Syntology 2026-04 (from id) mkboch/UDAE/evaluation/grader.py 87a574d68dd0da52 unverified no licence file found · pointer only
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning added by Syntology 2026-03 (from id) rujiewu/GDO/gdo/extract_six_metrics.py db008612f143341f ran · our draft was wrong fingerprinted MIT (permissive)
MPCEval: A Benchmark for Multi-Party Conversation Generation added by Syntology 2026-03 (from id) Owen-Yang-18/MPCEval/src/local_speaker/utils/common.py 490e489b13b16ab5 unverified no licence file found · pointer only
Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs added by Syntology 2026-01 (from id) bli1/tone-matters/VLM_Benchmark_GitHub_Ready/evaluation/score_hybrid_600_200.py c535a7f4683e3452 unverified MIT (permissive)
SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models added by Syntology 2026-01 (from id) Yunlin-Zeng/visual-podcast-VLM/evaluation_50_samples/evaluate_metrics.py 017711da98a034fb unverified licence not identified · pointer only
Transplant Then Regenerate: A New Paradigm for Text Data Augmentation added by Syntology 2025-08 (from id) 1024er/cbert_aug/text_classification/nlp_utils.py be8f967431da41d7 unverified no licence file found · pointer only
arXiv:2507.02834 2025-07 (from id) dhcode-cpp/X-R1/src/x_r1/rewards.py 11f2676b13fb9f69 unverified Apache-2.0 (permissive)
RedPajama: an Open Dataset for Training Large Language Models 19 Nov 2024 togethercomputer/redpajama-data/app/src/artifacts/utils/data_utils.py 0823f0fa0938dbfd unverified Apache-2.0 (permissive)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 4 Sep 2024 opendatalab/pm4bench/src/pm4bench/metrics.py abd00a80d4ff626d unverified Apache-2.0 (permissive)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model 7 Jun 2024 Edresson/ZS-TTS-Evaluation/utils/generic_utils.py 15b7e710b24d5a4c ran fingerprinted MIT (permissive)
KnowledgeHub: An end-to-end Tool for Assisted Scientific Discovery 16 May 2024 kermitt2/grobid/grobid-trainer/resources/dataset/segmentation/article/light/corpus/tei/analyze_notes.py 354d9000bf5db282 ran fingerprinted Apache-2.0 (permissive)
KnowledgeHub: An end-to-end Tool for Assisted Scientific Discovery 16 May 2024 kermitt2/grobid/grobid-trainer/resources/dataset/segmentation/article/light/corpus/tei/detailed_analysis.py b906150b79481c79 ran fingerprinted Apache-2.0 (permissive)
Faithful Logical Reasoning via Symbolic Chain-of-Thought 28 May 2024 Aiden0526/SymbCoT/baselines/evaluation.py 71986ed6ac7b9157 ran fingerprinted MIT (permissive)
PL-MTEB: Polish Massive Text Embedding Benchmark 16 May 2024 rafalposwiata/pl-mteb/tasks/preparation/cleaning.py d3d2208d6468ed08 ran fingerprinted no licence file found · pointer only
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach 24 Apr 2024 LoveCatc/supervised-llm-uncertainty-estimation/utils/data_entry.py e32f292f91eab5d8 ran fingerprinted no licence file found · pointer only
Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension 28 Feb 2024 fanyin3639/lid-hallucinationdetection/src/metrics.py 4e686c57f3e71d5f ran fingerprinted no licence file found · pointer only
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models 7 Jan 2024 google/speaker-id/DiarizationLM/diarizationlm/utils.py cca62276ec792e36 ran fingerprinted Apache-2.0 (permissive)
Language Generation from Brain Recordings 16 Nov 2023 yeziyi1998/brain-language-generation/language_generation/src/post_hoc_evaluate.py f04f39dee25fb155 ran · our draft was wrong fingerprinted licence not identified · pointer only
Implicit meta-learning may lead language models to trust more reliable sources 23 Oct 2023 krasheninnikov/internalization/src/metrics.py 12cb4aef8c35072c ran fingerprinted no licence file found · pointer only
Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference 4 Oct 2023 digitalharborfoundation/rag-for-math-qa/src/rag/retrieval.py 94af878e3d7b1853 ran fingerprinted MIT (permissive)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context 15 Sep 2023 k2-fsa/libriheavy/scripts/extract_and_normalize_transcript.py 5ddb3b78d7b37efe ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool 2023-08 (from id) isamu-isozaki/PentestGPT/pentestgpt/utils/chroma_vector_db.py edf2156bd952c7ed unverified MIT (permissive)
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning 20 May 2023 teacherpeterpan/logic-llm/models/evaluation.py 71986ed6ac7b9157 ran fingerprinted MIT (permissive)
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 vizuaraai/tiny-stories-regional/analysis/BERT_BLEU_eval.py 31f124b85a98aa6e unverified MIT (permissive)
Answering Questions by Meta-Reasoning over Multiple Chains of Thought 25 Apr 2023 oriyor/reasoning-on-cots/src/pred_evaluators/evaluation.py c30d6b506e3a5d6e ran · our draft was wrong fingerprinted MIT (permissive)
Continual Sequence Generation with Adaptive Compositional Modules 20 Mar 2022 GT-SALT/Adaptive-Compositional-Modules/metrics.py b971ea3b818d09b7 unverified MIT (permissive)
Attacking Open-domain Question Answering by Injecting Misinformation 15 Oct 2021 teacherpeterpan/contraqa/evaluate.py c30d6b506e3a5d6e ran · our draft was wrong fingerprinted MIT (permissive)
Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node Embeddings 21 Jun 2020 hugochan/IDGL/src/core/utils/eval_utils.py 38ffdcfda0ea6889 unverified Apache-2.0 (permissive)
Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus 27 Jan 2020 bangliu/ACS-QG/DA_main.py d53579a101181ef7 unverified MIT (permissive)
Libri-Light: A Benchmark for ASR with Limited or No Supervision 17 Dec 2019 identical code first harvested elsewhere 5ddb3b78d7b37efe ran · our draft was wrong fingerprinted licence of this copy not recorded
LAMOL: LAnguage MOdeling for Lifelong Language Learning 7 Sep 2019 jojotenya/LAMOL/metrics.py b971ea3b818d09b7 unverified MIT (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 autobotasia/vibert/pre-training-vibert.py b4d10db840d16868 unverified Apache-2.0 (permissive)
Graph Convolutional Networks for Text Classification 15 Sep 2018 koreyou/text-gcn-chainer/nlp_utils.py 2950d6b582fe43a8 unverified CC0-1.0 (permissive)
arXiv:openreview_XtIRCAEYoJ WayneTomas/Artemis/val/coco_detection/convert_to_coco_result.py fb5350be94c487a0 unverified Apache-2.0 (permissive)
arXiv:openreview_8ppVmLtA2V yanweiyue/Mem-T/llm_judge.py 988bb104b8b80d1e unverified Apache-2.0 (permissive)
arXiv:aaai_29863 facebookresearch/EditEval/src/preprocessing.py 223dafc9c2794bd1 unverified CC0-1.0 (permissive)
arXiv:2025.findings-emnlp.1015 fzp0424/MT-R1-Zero/eval/extract_to_eval.py 7715ccc1e32f7073 unverified Apache-2.0 (permissive)
arXiv:2025.findings-acl.712 cxcscmu/Craw4LLM/document_rater.py 8d60275b3f501f63 unverified MIT (permissive)
arXiv:2025.emnlp-industry.89 sssrlll/L4/metrics.py b971ea3b818d09b7 unverified MIT (permissive)
arXiv:2024.findings-emnlp.687 IAAR-Shanghai/FastMem/src/utils.py 15d3a384e09473fb unverified Apache-2.0 (permissive)
arXiv:2024.findings-eacl.35 naist-nlp/atd-mcl/src/util.py 77b47e3e5dacb774 unverified MIT (permissive)
arXiv:2024.acl-long.805 togethercomputer/RedPajama-Data/app/src/artifacts/utils/data_utils.py 0823f0fa0938dbfd unverified Apache-2.0 (permissive)
arXiv:2023.acl-long.736 felixgwu/FastFusionNet/qa/general_utils.py dd599b6afb9aad1b unverified MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections