Home › Code › exact_match_score

exact_match_score

Syntologyentry name in harvested coderead from the graph 2026-09-24

exact_match_score appears in the code Syntology harvested for 82 papers, as 35 distinct code bodies found in 87 places (a place is one code body under one paper). At least one of them ran in 65 of the papers; 17 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named exact_match_score do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 19 of the 35 distinct code bodies named exact_match_score; 16 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
4ran · violated contract
0ran · our draft was wrong
0ran · fixture could not drive it
15ran
16unverified
17fingerprinted

Licence is a property of each copy, so it is counted per place: 32 of the 87 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

82 papers shown of 82, newest first; 87 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 7 papers added by Syntology; 7 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology 2026-09 (from id) ZaiizaiZHANG/ISO-RAG/analyze_predictions.py 66804d06fa601b9c unverified no licence file found · pointer only
ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology 2026-09 (from id) ZaiizaiZHANG/ISO-RAG/analyze_retrieval_qa_transfer.py 209e457134cf7e77 unverified no licence file found · pointer only
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering added by Syntology 2026-04 (from id) zhuyjan/WikiSeeker/utils/infoseek_evaluation_utils.py 7055aa97f2bde50e unverified Apache-2.0 (permissive)
Temporal Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN added by Syntology 2026-03 (from id) terrierteam/temporal_conflicts/fact_mutability/analysis/f1_score.py fd5f06fcb775f9e0 unverified no licence file found · pointer only
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering added by Syntology 2026-02 (from id) How-Young-X/CompactRAG/src/metrics/F1Eval.py 1a90d2afb9055e47 unverified MIT (permissive)
Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval added by Syntology 7 Jan 2026 ChenWangHKU/DTR/evaluation/metrics.py 108d116eb5bdf952 unverified no licence file found · pointer only
BaseCal: Unsupervised Confidence Calibration via Base Model Signals added by Syntology 2026-01 (from id) Tan-Hexiang/BaseCal/gen_metric/em.py 8aa9ea4da99420bd unverified no licence file found · pointer only
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA added by Syntology 2025-10 (from id) SapienzaNLP/LiteraryQA/literaryqa/ngram_metrics.py b2e99f6b9eea4ae3 unverified no licence file found · pointer only
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? 26 Jun 2025 Haoang97/CausalProbe-2024/metrics.py d97ab73d08dfa3b9 unverified no licence file found · pointer only
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 22 May 2025 RUCAIBox/R1-Searcher/train/reward_server_qwen_zero.py 1332ed2c182add91 unverified MIT (permissive)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning 9 Dec 2024 SalesforceAIResearch/SiReRAG/evaluate_2wiki.py 9d0dc82a4491f803 ran · violated contract fingerprinted no licence file found · pointer only
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression 5 Oct 2024 eunseongc/r2c/LLM_inference.py 31f795a3e0c1be69 ran fingerprinted no licence file found · pointer only
Unlocking Continual Learning Abilities in Language Models 25 Jun 2024 wenyudu/migu/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
From Instance Training to Instruction Learning: Task Adapters Generation from Instructions 18 Jun 2024 Xnhyacinth/TAGI/src/compute_metrics.py e168ce39d04b76f1 ran Apache-2.0 (permissive)
Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning? 13 Jun 2024 zhaochen0110/cotempqa/config.py 1d5bd4c4648b1e9d ran fingerprinted no licence file found · pointer only
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching 10 Jun 2024 identical code first harvested elsewhere 2c66ac5f9374f1f7 ran · violated contract fingerprinted licence of this copy not recorded
BERTs are Generative In-Context Learners 7 Jun 2024 ltgoslo/bert-in-context/glue/record.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training 31 May 2024 calubkk/RAAT/tuner/metrics/em_f1.py 0e88ab2d4fdaca3c ran no licence file found · pointer only
Perturbation-Restrained Sequential Model Editing 27 May 2024 mjy1111/PRUNE/edit_load.py 4be72e172f90cb2c ran fingerprinted no licence file found · pointer only
Benchmarking Benchmark Leakage in Large Language Models 29 Apr 2024 gair-nlp/benbench/src/metric_utils.py 6828c333fed5ca6d ran fingerprinted no licence file found · pointer only
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency 18 Apr 2024 facebookresearch/multisense_consistency/utils/eval_metrics.py 8325d0eddc238131 ran fingerprinted licence not identified · pointer only
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation 28 Feb 2024 xsc1234/info-rag/training/train_info_rag.py 34916672e3ddaf0c ran fingerprinted no licence file found · pointer only
Analysing The Impact of Sequence Composition on Language Model Pre-Training 21 Feb 2024 yuzhaouoe/pretraining-data-packing/evaluation/eval_utils.py efa3643614e9c5f5 ran fingerprinted MIT (permissive)
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs 19 Feb 2024 plageon/slimplm/SKR/skr.py 0be284da9bf9ca86 ran fingerprinted no licence file found · pointer only
Discerning and Resolving Knowledge Conflicts through Adaptive Decoding with Contextual Information-Entropy Constraint 19 Feb 2024 identical code first harvested elsewhere bf75673211903e14 ran · violated contract fingerprinted licence of this copy not recorded
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering 4 Feb 2024 upper9527/gerea/leaderboard_evaluation.py 108d116eb5bdf952 unverified no licence file found · pointer only
Repeat After Me: Transformers are Better than State Space Models at Copying 1 Feb 2024 sjelassi/transformers_ssm_copy/pretrained_exps/qa_evaluation_utils.py 67b1c1bbca4316ab ran fingerprinted MIT (permissive)
SLANG: New Concept Comprehension of Large Language Models 23 Jan 2024 meirtz/focusonslang-toolbox/eval_urban_f1.py 978ce66ddf4e7b47 ran fingerprinted MIT (permissive)
TAP4LLM: Table Provider on Sampling, Augmenting, and Packing Semi-structured Data for Large Language Model Reasoning 14 Dec 2023 Y-Sui/GPT4Table/table_meets_llm/eval/evaluate_benchmark.py 79f3ecb504ea2b1e ran fingerprinted no licence file found · pointer only
Leveraging Structured Information for Explainable Multi-hop Question Answering and Reasoning 7 Nov 2023 bcdnlp/structure-qa/src/hotpot_evaluate.py 9d0dc82a4491f803 ran · violated contract fingerprinted GPL-3.0 (copyleft) · pointer only
Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks 1 Nov 2023 pluslabnlp/active-it/ActiveIT/src/compute_metrics.py e168ce39d04b76f1 ran no licence file found · pointer only
Orthogonal Subspace Learning for Language Model Continual Learning 22 Oct 2023 cmnfriend/o-lora/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection 17 Oct 2023 AkariAsai/self-rag/retrieval_lm/metrics.py d97ab73d08dfa3b9 unverified MIT (permissive)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration 14 Sep 2023 ukplab/catfood/src/shortcuts/evaluate.py e35cb1727750ade1 ran fingerprinted MIT (permissive)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration 14 Sep 2023 ukplab/catfood/src/calibration/baseline/calibration_metrics.py ee7ee45b82d4949f ran fingerprinted MIT (permissive)
$\rm SP^3$: Enhancing Structured Pruning via PCA Projection 31 Aug 2023 hyx1999/sp3/bert/metric/squad.py bf75673211903e14 ran · violated contract fingerprinted no licence file found · pointer only
Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection 17 Aug 2023 leezekun/instruction-following-robustness-eval/qa_utils.py bf75673211903e14 ran · violated contract fingerprinted no licence file found · pointer only
Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering 24 May 2023 aviclu/peekacross/trainer_seq2seq_qa.py 8fad2994edad57e5 unverified MIT (permissive)
Evaluating Open-QA Evaluation 21 May 2023 wangcunxiang/QA-Eval/lexical_match.py c4ae86237858e5d6 unverified Apache-2.0 (permissive)
Poisoning Language Models During Instruction Tuning 1 May 2023 alexwan0/poisoning-instruction-tuned-models/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
GPT-4 Technical Report 15 Mar 2023 AUCOHL/RTL-Repo/src/utils.py 2d1ad042f17b54ac unverified Apache-2.0 (permissive)
Investigating the Effectiveness of Task-Agnostic Prefix Prompt for Instruction Following 28 Feb 2023 seonghyeonye/icil/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? 23 Feb 2023 edchengg/infoseek_eval/infoseek_eval.py 7055aa97f2bde50e unverified MIT (permissive)
Robustness of Learning from Task Instructions 7 Dec 2022 jiashenggu/tk-instruct/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive Adaptation 19 Oct 2022 identical code first harvested elsewhere f6c275d6a18330a9 ran · violated contract fingerprinted licence of this copy not recorded
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks 16 Apr 2022 renzelou/pick-rank/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
Single-dataset Experts for Multi-dataset Question Answering 28 Sep 2021 princeton-nlp/MADE/src/utils/metrics.py bf75673211903e14 ran · violated contract fingerprinted MIT (permissive)
Phrase Retrieval Learns Passage Retrieval, Too 16 Sep 2021 princeton-nlp/DensePhrases/densephrases/utils/eval_utils.py 9d0dc82a4491f803 ran · violated contract fingerprinted Apache-2.0 (permissive)
ReasonBERT: Pre-trained to Reason with Distant Supervision 10 Sep 2021 sunlab-osu/reasonbert/model/metric.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Learning to Perturb Word Embeddings for Out-of-distribution QA 6 May 2021 seanie12/SWEP/mrqa_utils.py bf75673211903e14 ran · violated contract fingerprinted MIT (permissive)
AmbiFC: Fact-Checking Ambiguous Claims with Evidence 1 Apr 2021 knot-fit-but/claimdissector/src/common/eval_utils.py efa3643614e9c5f5 ran fingerprinted Apache-2.0 (permissive)
English Machine Reading Comprehension Datasets: A Survey 25 Jan 2021 identical code first harvested elsewhere 2c66ac5f9374f1f7 ran · violated contract fingerprinted licence of this copy not recorded
Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps 2 Nov 2020 Alab-NII/2wikimultihop/2wikimultihop_evaluate.py 9d0dc82a4491f803 ran · violated contract fingerprinted Apache-2.0 (permissive)
CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question Answering 30 Oct 2020 sunlab-osu/CliniQG4QA/QA/evaluate-v1.1_human_generated.py bf75673211903e14 ran · violated contract fingerprinted no licence file found · pointer only
Answering Open-Domain Questions of Varying Reasoning Steps from Text 23 Oct 2020 identical code first harvested elsewhere f6c275d6a18330a9 ran · violated contract fingerprinted licence of this copy not recorded
Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets 6 Aug 2020 facebookresearch/QA-Overlap/evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted no licence file found · pointer only
The Depth-to-Width Interplay in Self-Attention 22 Jun 2020 uf-hobi-informatics-lab/GatorTron/finetuning/qa/evaluate-v1.1.py bf75673211903e14 ran · violated contract fingerprinted MIT (permissive)
Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering 24 Nov 2019 AkariAsai/learning_to_retrieve_reasoning_paths/eval_utils.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
Contextualized Sparse Representations for Real-Time Open-Domain Question Answering 7 Nov 2019 jhyuklee/sparc/evaluate-v1.1.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Contextualized Sparse Representations for Real-Time Open-Domain Question Answering 7 Nov 2019 jhyuklee/sparc/eval_utils.py 9d0dc82a4491f803 ran · violated contract fingerprinted Apache-2.0 (permissive)
Do Multi-hop Readers Dream of Reasoning Chains? 31 Oct 2019 helloeve/bert-co-matching/evaluate-v1.1-original.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Do Multi-hop Readers Dream of Reasoning Chains? 31 Oct 2019 helloeve/bert-co-matching/evaluate-v1.1.py d299ec45effa875c unverified Apache-2.0 (permissive)
MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension 22 Oct 2019 mrqa/MRQA-Shared-Task-2019/mrqa_official_eval.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering 13 Sep 2019 ZhangShiyue/QGforQA/LIB/EVAL/evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
Self-Assembling Modular Networks for Interpretable Multi-Hop Reasoning 12 Sep 2019 jiangycTarheel/NMN-MultiHopQA/hotpotqa/evaluate-v1.1.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index 13 Jun 2019 uwnlp/denspi/evaluate-v1.1.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Neural Arabic Question Answering 12 Jun 2019 husseinmozannar/SOQAL/baselines_reading/evaluate_baselines.py e9b032891a5228a4 unverified MIT (permissive)
MultiQA: An Empirical Investigation of Generalization and Transfer in Reading Comprehension 31 May 2019 alontalmor/multiqa/common/official_eval.py f6c275d6a18330a9 ran · violated contract fingerprinted no licence file found · pointer only
Cognitive Graph for Multi-Hop Reading Comprehension at Scale 14 May 2019 THUDM/CogQA/hotpot_evaluate_v1.py 9d0dc82a4491f803 ran · violated contract fingerprinted MIT (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 Microsoft/AzureML-BERT/finetune/evaluate_squad.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering 25 Sep 2018 hotpotqa/hotpot/hotpot_evaluate_v1.py 9d0dc82a4491f803 ran · violated contract fingerprinted Apache-2.0 (permissive)
DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension 21 Apr 2018 duorc/duorc/evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
Phrase-Indexed Question Answering: A New Challenge for Scalable Document Comprehension 20 Apr 2018 uwnlp/piqa/squad/piqa_evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering 14 Nov 2017 shuohangwang/mprc/trainedmodel/evaluation/quasart/evaluate-v1.1.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering 14 Nov 2017 shuohangwang/mprc/trainedmodel/evaluation/unftriviaqa/triviaqa_evaluation.py 2c66ac5f9374f1f7 ran · violated contract fingerprinted Apache-2.0 (permissive)
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension 9 May 2017 mandarjoshi90/triviaqa/evaluation/triviaqa_evaluation.py 2c66ac5f9374f1f7 ran · violated contract fingerprinted Apache-2.0 (permissive)
Bidirectional Attention Flow for Machine Comprehension 5 Nov 2016 ghus75/Question_Answering/code/evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted Apache-2.0 (permissive)
Machine Comprehension Using Match-LSTM and Answer Pointer 29 Aug 2016 identical code first harvested elsewhere f6c275d6a18330a9 ran · violated contract fingerprinted licence of this copy not recorded
SQuAD: 100,000+ Questions for Machine Comprehension of Text 16 Jun 2016 identical code first harvested elsewhere f6c275d6a18330a9 ran · violated contract fingerprinted licence of this copy not recorded
Character-Aware Neural Language Models 26 Aug 2015 NLPLearn/QANet/evaluate-v1.1.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
arXiv:aaai_34690 ZongyueQin/DSBD/sampling/utils.py bf75673211903e14 ran · violated contract fingerprinted MIT (permissive)
arXiv:2025.acl-long.191 UKPLab/acl2025-diverse-cot/src/hotpotqa_evaluation.py 9d0dc82a4491f803 ran · violated contract fingerprinted Apache-2.0 (permissive)
arXiv:2024.findings-emnlp.379 wenyudu/MIGU/src/compute_metrics.py e168ce39d04b76f1 ran MIT (permissive)
arXiv:2023.findings-emnlp.836 IBM/ensemble-instruct/ensemble_instruct/ensemble_output.py 3c8cf46311df61a2 unverified Apache-2.0 (permissive)
arXiv:2023.findings-emnlp.835 THU-KEG/ProbTree/src/2wiki/RoHT/evaluate.py 9d0dc82a4491f803 ran · violated contract fingerprinted MIT (permissive)
arXiv:2023.emnlp-main.803 yisunlp/Anti-CF/utils/squad_evaluate.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)
arXiv:2021.acl-long.48 PluviophileYU/COSY/XQA/src/evaluate_v1_1.py f6c275d6a18330a9 ran · violated contract fingerprinted MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections