Home › Code › count_words

count_words

Syntologyentry name in harvested coderead from the graph 2026-09-24

count_words appears in the code Syntology harvested for 28 papers, as 21 distinct code bodies found in 28 places (a place is one code body under one paper). At least one of them ran in 17 of the papers; 11 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named count_words do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 11 of the 21 distinct code bodies named count_words; 10 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

3ran · honoured contract
0ran · violated contract
0ran · our draft was wrong
0ran · fixture could not drive it
8ran
10unverified
11fingerprinted

Licence is a property of each copy, so it is counted per place: 11 of the 28 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

28 papers shown of 28, newest first; 28 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 7 papers added by Syntology; 1 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text added by Syntology 2026-08 (from id) trieuntu/VietAIDetector/preprocessing/text_utils.py 510c69b7c1c15674 ran fingerprinted MIT (permissive)
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia added by Syntology 2026-06 (from id) cepin19/CzechDocs/count_curated_stats.py a379d71b3bd8ea58 ran fingerprinted no licence file found · pointer only
Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen added by Syntology 2026-04 (from id) synthiumjp/koriat/build_cues.py 3ede3ad4d2b8d09f ran fingerprinted MIT (permissive)
Can LLMs Understand the Impact of Trauma? Costs and Benefits of LLMs Coding the Interviews of Firearm Violence Survivors added by Syntology 2026-04 (from id) jhzsquared/AIvsHumanCoding/count_data.py 9a0e57cbbb4a12e8 unverified no licence file found · pointer only
CLASE: A Hybrid Method for Chinese Legalese Stylistic Evaluation added by Syntology 2026-02 (from id) rexera/CLASE/linguistic_features/feature_extractor.py 9c81670a59d88906 ran · honoured contract fingerprinted MIT (permissive)
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models added by Syntology 2025-10 (from id) thierry123454/tool-selection-bias/5_bias_investigation/continued_pretraining/generate_data_gemini.py 2498e76b5f8a7bd9 unverified no licence file found · pointer only
Transplant Then Regenerate: A New Paradigm for Text Data Augmentation added by Syntology 2025-08 (from id) 1024er/cbert_aug/utils.py ee80d8c65de13c06 unverified no licence file found · pointer only
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning 23 Jun 2025 thudm/longwriter/evaluation/pred.py 2bb5c6316f656bec unverified Apache-2.0 (permissive)
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions 21 Nov 2024 aidc-ai/marco-o1/src/v2/src/tree_search/evaluator/ifeval.py d23b2037805b8d2b unverified licence not identified · pointer only
Language Models can Self-Lengthen to Generate Long Texts 31 Oct 2024 QwenLM/Self-Lengthen/eval/length_following_eval.py b3334d24e49ff2b5 ran · honoured contract fingerprinted Apache-2.0 (permissive)
Back to School: Translation Using Grammar Books 20 Oct 2024 jonathanhus/back-to-school/utilities/count_data.py bdd4c45f6daf0dab unverified MIT (permissive)
Meta-Chunking: Learning Text Segmentation and Semantic Completion via Logical Perception 16 Oct 2024 IAAR-Shanghai/Meta-Chunking/meta_chunking/LongBench/LumberChunker.py a4d6434de692eabe ran fingerprinted Apache-2.0 (permissive)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation 12 Oct 2024 dongguanting/FollowRAG/FollowRAG/utils/instruction_following_eval/instructions_util.py cdcc85ca09b00f7e ran fingerprinted no licence file found · pointer only
The Rise of AI-Generated Content in Wikipedia 10 Oct 2024 brooksca3/wiki_collection/misc/get_hyperlinks.py 95141547985bc194 ran fingerprinted no licence file found · pointer only
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models 4 Aug 2024 wbw520/DiReCT/statistics.py 8c024efc9f977616 ran fingerprinted no licence file found · pointer only
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches 1 Jul 2024 henryzhongsc/longctx_bench/eval/passkey_utils/passkey_utils.py 2207706bbe348d2c ran fingerprinted MIT (permissive)
LumberChunker: Long-Form Narrative Document Segmentation 25 Jun 2024 joaodsmarques/lumberchunker/Code/LumberChunker-Segmentation.py 70c76dd00047fdbf ran · honoured contract fingerprinted no licence file found · pointer only
Native Design Bias: Studying the Impact of English Nativeness on Language Model Performance 25 Jun 2024 manon-reusens/native_en_bias/dataset_statistics.py 0b33a7e746e505d2 unverified no licence file found · pointer only
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones? 18 Jun 2024 QwenLM/ConsisEval/instruction_following_check/instructions_util.py cdcc85ca09b00f7e ran fingerprinted MIT (permissive)
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models 24 Apr 2024 meowpass/followcomplexinstruction/get_data/instructions_util.py cdcc85ca09b00f7e ran fingerprinted no licence file found · pointer only
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 kwanwaichung/mt-eval/utils/global_inst.py cdcc85ca09b00f7e ran fingerprinted MIT (permissive)
Instruction-Following Evaluation for Large Language Models 14 Nov 2023 josejg/instruction_following_eval/instruction_following_eval/instructions_util.py cdcc85ca09b00f7e ran fingerprinted Apache-2.0 (permissive)
YaRN: Efficient Context Window Extension of Large Language Models 31 Aug 2023 qwenlm/qwq/eval/eval/ifeval_utils/instructions_util.py cdcc85ca09b00f7e ran fingerprinted Apache-2.0 (permissive)
Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes 20 Aug 2020 allenai/scruples/src/scruples/utils.py d6c73ac35e8781c2 unverified Apache-2.0 (permissive)
Graph Universal Adversarial Attacks: A Few Bad Actors Ruin Graph Learning Models 12 Feb 2020 chisam0217/Graph-Universal-Attack/deepwalk/evaluate_deepwalk.py b30d29f66472e711 unverified MIT (permissive)
RWR-GAE: Random Walk Regularization for Graph Auto Encoders 12 Aug 2019 MysteryVaibhav/DW-GAE/deepWalk/walks.py a236ee68c305670f unverified MIT (permissive)
Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations 16 May 2018 pfnet-research/contextual_augmentation/utils.py ee80d8c65de13c06 unverified MIT (permissive)
arXiv:2025.emnlp-main.1010 joaopfonseca/SafeNudge/ctg/instructions_util.py cdcc85ca09b00f7e ran fingerprinted MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections