Home › Code › normalize_word

normalize_word

Syntologyentry name in harvested coderead from the graph 2026-09-24

normalize_word appears in the code Syntology harvested for 22 papers, as 9 distinct code bodies found in 24 places (a place is one code body under one paper). At least one of them ran in 4 of the papers; 2 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named normalize_word do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 2 of the 9 distinct code bodies named normalize_word; 7 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
1ran · our draft was wrong
0ran · fixture could not drive it
1ran
7unverified
2fingerprinted

Licence is a property of each copy, so it is counted per place: 4 of the 24 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

22 papers shown of 22, newest first; 24 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 1 papers added by Syntology; 2 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific Literature added by Syntology 2026-03 (from id) basiralab/HGNet/RE/baselines/PL-Marker/preprocess_ontonotes.py ba4eb9ed094dc59e unverified MIT (permissive)
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs 27 Mar 2025 CVI-SZU/FaceBench/evaluation/utils.py 8f1a44ab20956301 unverified MIT (permissive)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization 9 Dec 2024 aiming-lab/mmedpo/eval/eval_report.py 8f1a44ab20956301 unverified Apache-2.0 (permissive)
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases 3 Dec 2024 kki2eve/agri-llava/agri_llava/eval/eval_metrics/glossary.py 8f1a44ab20956301 unverified Apache-2.0 (permissive)
EMR-Merging: Tuning-Free High-Performance Model Merging 23 May 2024 harveyhuang18/emr_merging/merge_beit3/glossary.py 8f1a44ab20956301 unverified no licence file found · pointer only
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering 19 Apr 2024 garygutc/lapa_model/prepro/glossary.py 8f1a44ab20956301 unverified no licence file found · pointer only
Seq2seq is All You Need for Coreference Resolution 20 Oct 2023 WenzhengZhang/Seq2seqCoref/data.py 1dfbebba563fe3db ran fingerprinted MIT (permissive)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs 1 Oct 2023 sy-xuan/pink/pink/eval/model_vqav2.py 8f1a44ab20956301 unverified no licence file found · pointer only
ArtWhisperer: A Dataset for Characterizing Human-AI Interactions in Artistic Creations 13 Jun 2023 kailas-v/ArtWhisperer/src/analysis.py 5e60c034ab793b02 unverified MIT (permissive)
Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners 18 May 2023 eric-ai-lab/dsd/utils/glossary.py 8f1a44ab20956301 unverified MIT (permissive)
Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role Interactions 26 May 2022 AIPHES/emnlp19-moverscore/webservice/server/server/metrics/utils.py 5991dab5f3e534fe unverified MIT (permissive)
LingMess: Linguistically Informed Multi Expert Scorers for Coreference Resolution 25 May 2022 shon-otmazgin/lingmess-coref/prepare_ontonotes/minimize.py 3b625164c1147fcf unverified MIT (permissive)
Incorporating Constituent Syntax for Coreference Resolution 22 Feb 2022 mandarjoshi90/coref/minimize.py 3b625164c1147fcf unverified Apache-2.0 (permissive)
Packed Levitated Marker for Entity and Relation Extraction 13 Sep 2021 thunlp/pl-marker/preprocess_ontonotes.py ba4eb9ed094dc59e unverified MIT (permissive)
DWIE: an entity-centric dataset for multi-task document-level information extraction 26 Sep 2020 klimzaporojets/consistent-el/src/data_processing/bert_preprocessing.py 67b9b442fafaf135 ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
Hierarchical Contextualized Representation for Named Entity Recognition 6 Nov 2019 cslydia/Hire-NER/utils/functions.py 1b6406aa4e398f9e unverified Apache-2.0 (permissive)
SpanBERT: Improving Pre-training by Representing and Predicting Spans 24 Jul 2019 amore-upf/masked-coreference/minimize.py 3b625164c1147fcf unverified Apache-2.0 (permissive)
NCRF++: An Open-source Neural Sequence Labeling Toolkit 14 Jun 2018 jiesutd/PyTorchSeqLabel/utils/functions.py 1b6406aa4e398f9e unverified Apache-2.0 (permissive)
Higher-order Coreference Resolution with Coarse-to-fine Inference 15 Apr 2018 Filter-Bubble/e2e-Dutch/e2edutch/minimize.py 67b9b442fafaf135 ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
Higher-order Coreference Resolution with Coarse-to-fine Inference 15 Apr 2018 kentonl/e2e-coref/minimize.py 3b625164c1147fcf unverified Apache-2.0 (permissive)
End-to-end Neural Coreference Resolution 21 Jul 2017 Achint08/e2e-coref-keras/preprop.py 67b9b442fafaf135 ran · our draft was wrong fingerprinted no licence file found · pointer only
arXiv:2025.findings-emnlp.1354 basiralab/DP-ERE/baselines/PL-Marker/preprocess_ontonotes.py ba4eb9ed094dc59e unverified MIT (permissive)
arXiv:2021.naacl-main.125 Fantabulous-J/coref-HGAT/extract_constituency.py 3b03a04bf9022a1c unverified Apache-2.0 (permissive)
arXiv:2021.naacl-main.125 Fantabulous-J/coref-HGAT/minimize.py 3b625164c1147fcf unverified Apache-2.0 (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections