| ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology |
2026-09 (from id) |
ZaiizaiZHANG/ISO-RAG/analyze_predictions.py 66804d06fa601b9c |
unverified |
no licence file found · pointer only |
| ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology |
2026-09 (from id) |
ZaiizaiZHANG/ISO-RAG/analyze_retrieval_qa_transfer.py 209e457134cf7e77 |
unverified |
no licence file found · pointer only |
| WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering added by Syntology |
2026-04 (from id) |
zhuyjan/WikiSeeker/utils/infoseek_evaluation_utils.py 7055aa97f2bde50e |
unverified |
Apache-2.0 (permissive) |
| Temporal Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN added by Syntology |
2026-03 (from id) |
terrierteam/temporal_conflicts/fact_mutability/analysis/f1_score.py fd5f06fcb775f9e0 |
unverified |
no licence file found · pointer only |
| CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering added by Syntology |
2026-02 (from id) |
How-Young-X/CompactRAG/src/metrics/F1Eval.py 1a90d2afb9055e47 |
unverified |
MIT (permissive) |
| Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval added by Syntology |
7 Jan 2026 |
ChenWangHKU/DTR/evaluation/metrics.py 108d116eb5bdf952 |
unverified |
no licence file found · pointer only |
| BaseCal: Unsupervised Confidence Calibration via Base Model Signals added by Syntology |
2026-01 (from id) |
Tan-Hexiang/BaseCal/gen_metric/em.py 8aa9ea4da99420bd |
unverified |
no licence file found · pointer only |
| LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA added by Syntology |
2025-10 (from id) |
SapienzaNLP/LiteraryQA/literaryqa/ngram_metrics.py b2e99f6b9eea4ae3 |
unverified |
no licence file found · pointer only |
| Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? |
26 Jun 2025 |
Haoang97/CausalProbe-2024/metrics.py d97ab73d08dfa3b9 |
unverified |
no licence file found · pointer only |
| R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning |
22 May 2025 |
RUCAIBox/R1-Searcher/train/reward_server_qwen_zero.py 1332ed2c182add91 |
unverified |
MIT (permissive) |
| SiReRAG: Indexing Similar and Related Information for Multihop Reasoning |
9 Dec 2024 |
SalesforceAIResearch/SiReRAG/evaluate_2wiki.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression |
5 Oct 2024 |
eunseongc/r2c/LLM_inference.py 31f795a3e0c1be69 |
ran
fingerprinted |
no licence file found · pointer only |
| Unlocking Continual Learning Abilities in Language Models |
25 Jun 2024 |
wenyudu/migu/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| From Instance Training to Instruction Learning: Task Adapters Generation from Instructions |
18 Jun 2024 |
Xnhyacinth/TAGI/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
Apache-2.0 (permissive) |
| Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning? |
13 Jun 2024 |
zhaochen0110/cotempqa/config.py 1d5bd4c4648b1e9d |
ran
fingerprinted |
no licence file found · pointer only |
| Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching |
10 Jun 2024 |
identical code first harvested elsewhere 2c66ac5f9374f1f7 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| BERTs are Generative In-Context Learners |
7 Jun 2024 |
ltgoslo/bert-in-context/glue/record.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training |
31 May 2024 |
calubkk/RAAT/tuner/metrics/em_f1.py 0e88ab2d4fdaca3c |
ran
|
no licence file found · pointer only |
| Perturbation-Restrained Sequential Model Editing |
27 May 2024 |
mjy1111/PRUNE/edit_load.py 4be72e172f90cb2c |
ran
fingerprinted |
no licence file found · pointer only |
| Benchmarking Benchmark Leakage in Large Language Models |
29 Apr 2024 |
gair-nlp/benbench/src/metric_utils.py 6828c333fed5ca6d |
ran
fingerprinted |
no licence file found · pointer only |
| From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency |
18 Apr 2024 |
facebookresearch/multisense_consistency/utils/eval_metrics.py 8325d0eddc238131 |
ran
fingerprinted |
licence not identified · pointer only |
| Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation |
28 Feb 2024 |
xsc1234/info-rag/training/train_info_rag.py 34916672e3ddaf0c |
ran
fingerprinted |
no licence file found · pointer only |
| Analysing The Impact of Sequence Composition on Language Model Pre-Training |
21 Feb 2024 |
yuzhaouoe/pretraining-data-packing/evaluation/eval_utils.py efa3643614e9c5f5 |
ran
fingerprinted |
MIT (permissive) |
| Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs |
19 Feb 2024 |
plageon/slimplm/SKR/skr.py 0be284da9bf9ca86 |
ran
fingerprinted |
no licence file found · pointer only |
| Discerning and Resolving Knowledge Conflicts through Adaptive Decoding with Contextual Information-Entropy Constraint |
19 Feb 2024 |
identical code first harvested elsewhere bf75673211903e14 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering |
4 Feb 2024 |
upper9527/gerea/leaderboard_evaluation.py 108d116eb5bdf952 |
unverified |
no licence file found · pointer only |
| Repeat After Me: Transformers are Better than State Space Models at Copying |
1 Feb 2024 |
sjelassi/transformers_ssm_copy/pretrained_exps/qa_evaluation_utils.py 67b1c1bbca4316ab |
ran
fingerprinted |
MIT (permissive) |
| SLANG: New Concept Comprehension of Large Language Models |
23 Jan 2024 |
meirtz/focusonslang-toolbox/eval_urban_f1.py 978ce66ddf4e7b47 |
ran
fingerprinted |
MIT (permissive) |
| TAP4LLM: Table Provider on Sampling, Augmenting, and Packing Semi-structured Data for Large Language Model Reasoning |
14 Dec 2023 |
Y-Sui/GPT4Table/table_meets_llm/eval/evaluate_benchmark.py 79f3ecb504ea2b1e |
ran
fingerprinted |
no licence file found · pointer only |
| Leveraging Structured Information for Explainable Multi-hop Question Answering and Reasoning |
7 Nov 2023 |
bcdnlp/structure-qa/src/hotpot_evaluate.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
GPL-3.0 (copyleft) · pointer only |
| Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks |
1 Nov 2023 |
pluslabnlp/active-it/ActiveIT/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
no licence file found · pointer only |
| Orthogonal Subspace Learning for Language Model Continual Learning |
22 Oct 2023 |
cmnfriend/o-lora/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection |
17 Oct 2023 |
AkariAsai/self-rag/retrieval_lm/metrics.py d97ab73d08dfa3b9 |
unverified |
MIT (permissive) |
| CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration |
14 Sep 2023 |
ukplab/catfood/src/shortcuts/evaluate.py e35cb1727750ade1 |
ran
fingerprinted |
MIT (permissive) |
| CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration |
14 Sep 2023 |
ukplab/catfood/src/calibration/baseline/calibration_metrics.py ee7ee45b82d4949f |
ran
fingerprinted |
MIT (permissive) |
| $\rm SP^3$: Enhancing Structured Pruning via PCA Projection |
31 Aug 2023 |
hyx1999/sp3/bert/metric/squad.py bf75673211903e14 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection |
17 Aug 2023 |
leezekun/instruction-following-robustness-eval/qa_utils.py bf75673211903e14 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering |
24 May 2023 |
aviclu/peekacross/trainer_seq2seq_qa.py 8fad2994edad57e5 |
unverified |
MIT (permissive) |
| Evaluating Open-QA Evaluation |
21 May 2023 |
wangcunxiang/QA-Eval/lexical_match.py c4ae86237858e5d6 |
unverified |
Apache-2.0 (permissive) |
| Poisoning Language Models During Instruction Tuning |
1 May 2023 |
alexwan0/poisoning-instruction-tuned-models/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| GPT-4 Technical Report |
15 Mar 2023 |
AUCOHL/RTL-Repo/src/utils.py 2d1ad042f17b54ac |
unverified |
Apache-2.0 (permissive) |
| Investigating the Effectiveness of Task-Agnostic Prefix Prompt for Instruction Following |
28 Feb 2023 |
seonghyeonye/icil/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? |
23 Feb 2023 |
edchengg/infoseek_eval/infoseek_eval.py 7055aa97f2bde50e |
unverified |
MIT (permissive) |
| Robustness of Learning from Task Instructions |
7 Dec 2022 |
jiashenggu/tk-instruct/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive Adaptation |
19 Oct 2022 |
identical code first harvested elsewhere f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks |
16 Apr 2022 |
renzelou/pick-rank/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| Single-dataset Experts for Multi-dataset Question Answering |
28 Sep 2021 |
princeton-nlp/MADE/src/utils/metrics.py bf75673211903e14 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Phrase Retrieval Learns Passage Retrieval, Too |
16 Sep 2021 |
princeton-nlp/DensePhrases/densephrases/utils/eval_utils.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| ReasonBERT: Pre-trained to Reason with Distant Supervision |
10 Sep 2021 |
sunlab-osu/reasonbert/model/metric.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Learning to Perturb Word Embeddings for Out-of-distribution QA |
6 May 2021 |
seanie12/SWEP/mrqa_utils.py bf75673211903e14 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| AmbiFC: Fact-Checking Ambiguous Claims with Evidence |
1 Apr 2021 |
knot-fit-but/claimdissector/src/common/eval_utils.py efa3643614e9c5f5 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| English Machine Reading Comprehension Datasets: A Survey |
25 Jan 2021 |
identical code first harvested elsewhere 2c66ac5f9374f1f7 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps |
2 Nov 2020 |
Alab-NII/2wikimultihop/2wikimultihop_evaluate.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question Answering |
30 Oct 2020 |
sunlab-osu/CliniQG4QA/QA/evaluate-v1.1_human_generated.py bf75673211903e14 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| Answering Open-Domain Questions of Varying Reasoning Steps from Text |
23 Oct 2020 |
identical code first harvested elsewhere f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets |
6 Aug 2020 |
facebookresearch/QA-Overlap/evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| The Depth-to-Width Interplay in Self-Attention |
22 Jun 2020 |
uf-hobi-informatics-lab/GatorTron/finetuning/qa/evaluate-v1.1.py bf75673211903e14 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering |
24 Nov 2019 |
AkariAsai/learning_to_retrieve_reasoning_paths/eval_utils.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Contextualized Sparse Representations for Real-Time Open-Domain Question Answering |
7 Nov 2019 |
jhyuklee/sparc/evaluate-v1.1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Contextualized Sparse Representations for Real-Time Open-Domain Question Answering |
7 Nov 2019 |
jhyuklee/sparc/eval_utils.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Do Multi-hop Readers Dream of Reasoning Chains? |
31 Oct 2019 |
helloeve/bert-co-matching/evaluate-v1.1-original.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Do Multi-hop Readers Dream of Reasoning Chains? |
31 Oct 2019 |
helloeve/bert-co-matching/evaluate-v1.1.py d299ec45effa875c |
unverified |
Apache-2.0 (permissive) |
| MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension |
22 Oct 2019 |
mrqa/MRQA-Shared-Task-2019/mrqa_official_eval.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering |
13 Sep 2019 |
ZhangShiyue/QGforQA/LIB/EVAL/evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Self-Assembling Modular Networks for Interpretable Multi-Hop Reasoning |
12 Sep 2019 |
jiangycTarheel/NMN-MultiHopQA/hotpotqa/evaluate-v1.1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index |
13 Jun 2019 |
uwnlp/denspi/evaluate-v1.1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Neural Arabic Question Answering |
12 Jun 2019 |
husseinmozannar/SOQAL/baselines_reading/evaluate_baselines.py e9b032891a5228a4 |
unverified |
MIT (permissive) |
| MultiQA: An Empirical Investigation of Generalization and Transfer in Reading Comprehension |
31 May 2019 |
alontalmor/multiqa/common/official_eval.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| Cognitive Graph for Multi-Hop Reading Comprehension at Scale |
14 May 2019 |
THUDM/CogQA/hotpot_evaluate_v1.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding |
11 Oct 2018 |
Microsoft/AzureML-BERT/finetune/evaluate_squad.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering |
25 Sep 2018 |
hotpotqa/hotpot/hotpot_evaluate_v1.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension |
21 Apr 2018 |
duorc/duorc/evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| Phrase-Indexed Question Answering: A New Challenge for Scalable Document Comprehension |
20 Apr 2018 |
uwnlp/piqa/squad/piqa_evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering |
14 Nov 2017 |
shuohangwang/mprc/trainedmodel/evaluation/quasart/evaluate-v1.1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering |
14 Nov 2017 |
shuohangwang/mprc/trainedmodel/evaluation/unftriviaqa/triviaqa_evaluation.py 2c66ac5f9374f1f7 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension |
9 May 2017 |
mandarjoshi90/triviaqa/evaluation/triviaqa_evaluation.py 2c66ac5f9374f1f7 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Bidirectional Attention Flow for Machine Comprehension |
5 Nov 2016 |
ghus75/Question_Answering/code/evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Machine Comprehension Using Match-LSTM and Answer Pointer |
29 Aug 2016 |
identical code first harvested elsewhere f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| SQuAD: 100,000+ Questions for Machine Comprehension of Text |
16 Jun 2016 |
identical code first harvested elsewhere f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
licence of this copy not recorded |
| Character-Aware Neural Language Models |
26 Aug 2015 |
NLPLearn/QANet/evaluate-v1.1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| arXiv:aaai_34690 |
|
ZongyueQin/DSBD/sampling/utils.py bf75673211903e14 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| arXiv:2025.acl-long.191 |
|
UKPLab/acl2025-diverse-cot/src/hotpotqa_evaluation.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.379 |
|
wenyudu/MIGU/src/compute_metrics.py e168ce39d04b76f1 |
ran
|
MIT (permissive) |
| arXiv:2023.findings-emnlp.836 |
|
IBM/ensemble-instruct/ensemble_instruct/ensemble_output.py 3c8cf46311df61a2 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2023.findings-emnlp.835 |
|
THU-KEG/ProbTree/src/2wiki/RoHT/evaluate.py 9d0dc82a4491f803 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| arXiv:2023.emnlp-main.803 |
|
yisunlp/Anti-CF/utils/squad_evaluate.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| arXiv:2021.acl-long.48 |
|
PluviophileYU/COSY/XQA/src/evaluate_v1_1.py f6c275d6a18330a9 |
ran · violated contract
fingerprinted |
MIT (permissive) |