| SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification added by Syntology |
2026-09 (from id) |
baranwa2/SOVER/code/Variablemapping_extraction.py f78f6711b1fdcd4f |
unverified |
licence not identified · pointer only |
| Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning added by Syntology |
2026-07 (from id) |
aixiuxiuxiu/Legal-MT-SFT-RL/convert_jsonl_to_prompts.py 07986447a705f50c |
ran
|
MIT (permissive) |
| Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation added by Syntology |
2026-06 (from id) |
larchlab/Illusions-of-the-Gold-Standard/llm_label_v5.py 6f867f364d45edd5 |
ran
fingerprinted |
MIT (permissive) |
| A Case Study on the Impact of Anonymization Along the RAG Pipeline added by Syntology |
2026-04 (from id) |
andreea-bodea/GuardRAG/src/Presidio/Presidio_OpenAI.py 3814d2a46da525ad |
unverified |
no licence file found · pointer only |
| A Case Study on the Impact of Anonymization Along the RAG Pipeline added by Syntology |
2026-04 (from id) |
andreea-bodea/GuardRAG/src/Presidio/Presidio_OpenAI.py 648834d74ca15c46 |
unverified |
no licence file found · pointer only |
| Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning added by Syntology |
2026-02 (from id) |
wdqqdw/Echo/inference/inference_multiturn.py 5def991b7b2b440d |
unverified |
no licence file found · pointer only |
| BaseCal: Unsupervised Confidence Calibration via Base Model Signals added by Syntology |
2026-01 (from id) |
Tan-Hexiang/BaseCal/src/prompts.py db9655767f69c29f |
unverified |
no licence file found · pointer only |
| BaseCal: Unsupervised Confidence Calibration via Base Model Signals added by Syntology |
2026-01 (from id) |
Tan-Hexiang/BaseCal/src/prompts_wo_prefix_answer.py d16a3b512854596e |
unverified |
no licence file found · pointer only |
| CORRECT: Condensed Error Recognition via Knowledge Transfer in Multi-agent Systems added by Syntology |
2025-09 (from id) |
UIUC-MLSys/CORRECT/src/error_schema_generator_cloud.py 4e0a44a338e3f88e |
unverified |
licence not identified · pointer only |
| Effective Training Data Synthesis for Improving MLLM Chart Understanding added by Syntology |
2025-08 (from id) |
yuweiyang-anu/ECD/data_generation_pipeline/figure_size_post_processing.py b740188ede016838 |
unverified |
MIT (permissive) |
| A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning |
11 Jul 2025 |
analokmaus/kaggle-aimo2-fast-math-r1/experiments/train_fast_nemotron_14b.py 522a7f189e5ea3b4 |
unverified |
no licence file found · pointer only |
| A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning |
11 Jul 2025 |
analokmaus/kaggle-aimo2-fast-math-r1/experiments/train_first_stage.py e14b5e417be024bf |
unverified |
no licence file found · pointer only |
| dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching |
17 May 2025 |
maomaocun/dLLM-cache/LLama_test_flops_script.py 0976b0b29f8d1d22 |
unverified |
Apache-2.0 (permissive) |
| Social Sycophancy: A Broader Understanding of LLM Sycophancy |
20 May 2025 |
myracheng/elephant/sycophancy_scorers.py 01384b0681810072 |
ran · our draft was wrong
|
CC0-1.0 (permissive) |
| Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning |
8 Apr 2025 |
Sanchit-404/Reasoing-Towards-Fairness/scripts/extract_reasoning_traces.py f2ba52311094f608 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Political Neutrality in AI Is Impossible- But Here Is How to Approximate It |
18 Feb 2025 |
jfisher52/approximation_political_neutrality/src/eval_templates.py 29e163b5d3fd0c27 |
ran · our draft was wrong
|
GPL-3.0 (copyleft) · pointer only |
| LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content |
14 Oct 2024 |
nimrodshabtay/livexiv/vqa_generation/filter_utils.py 730b80d67018d2aa |
ran
|
Apache-2.0 (permissive) |
| Do Unlearning Methods Remove Information from Language Model Weights? |
11 Oct 2024 |
aghyad-deeb/unlearning_evaluation/unlearn_corpus.py 814f2de8f4bd281e |
ran · our draft was wrong
|
no licence file found · pointer only |
| Do Unlearning Methods Remove Information from Language Model Weights? |
11 Oct 2024 |
aghyad-deeb/unlearning_evaluation/finetune_corpus.py 246c7ef8f90b3e7e |
ran · our draft was wrong
|
no licence file found · pointer only |
| Mitigating the Language Mismatch and Repetition Issues in LLM-based Machine Translation via Model Editing |
9 Oct 2024 |
weichuanw/llm-based-mt-via-model-editing/Model_Editing_Adaptation/fv_task_adaptation/src/utils/prompt_utils.py 5daf29ec78f21d13 |
ran
|
no licence file found · pointer only |
| Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought |
8 Oct 2024 |
LightChen233/reasoning-boundary/request_text.py 46be0dbb921f8d85 |
ran
|
no licence file found · pointer only |
| Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought |
8 Oct 2024 |
LightChen233/reasoning-boundary/request_multimodal.py 419ea57087ee3539 |
unverified |
no licence file found · pointer only |
| Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions |
3 Oct 2024 |
MichiganNLP/MultiAgent_ImplicitBias/Code/mistral_ft.py a2ba955201953f73 |
ran
|
no licence file found · pointer only |
| CompAct: Compressing Retrieved Documents Actively for Question Answering |
12 Jul 2024 |
dmis-lab/CompAct/utils.py de1f2a7546b67b6a |
ran
|
no licence file found · pointer only |
| The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models |
28 Jun 2024 |
shin-ee-chen/SIFo/llm_inference/get_responses_from_vllm_models.py f2a96da09d0ad98f |
ran · our draft was wrong
|
no licence file found · pointer only |
| AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries |
27 Jun 2024 |
saparina/ambrosia/src/db_generation/generate_databases.py ba28dc7e25b4ea7b |
unverified |
no licence file found · pointer only |
| Evidence of a log scaling law for political persuasion with large language models |
20 Jun 2024 |
kobihackenburg/scaling-llm-persuasion/main_study/code/01_instructionTune.py 40357b3d64db3d0d |
ran
|
MIT (permissive) |
| MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction |
18 Jun 2024 |
nyushcs/moleculargpt/ICL_test_diversity.py 5cbb7bb23f2c061d |
ran
|
Apache-2.0 (permissive) |
| MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction |
18 Jun 2024 |
nyushcs/moleculargpt/ICL_test_reverse_cls.py 437f17c491f15275 |
ran
|
Apache-2.0 (permissive) |
| MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction |
18 Jun 2024 |
nyushcs/moleculargpt/ICL_test_reverse_reg.py 40ec3f3c72f787b8 |
ran
|
Apache-2.0 (permissive) |
| MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction |
18 Jun 2024 |
nyushcs/moleculargpt/ICL_test_sim_cls.py 504b9ab37dc7458f |
ran
|
Apache-2.0 (permissive) |
| MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction |
18 Jun 2024 |
nyushcs/moleculargpt/ICL_test_sim_reg.py 32661c96034aa972 |
ran
|
Apache-2.0 (permissive) |
| MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding |
13 Jun 2024 |
muirbench/MuirBench/eval/utils/preprocess.py 1cc5bed6d9e89200 |
ran · our draft was wrong
|
no licence file found · pointer only |
| CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery |
12 Jun 2024 |
csbench/csbench/vllm-main/examples/csbench/gen_model_answer_en.py 05cde4d9afed5c53 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models |
11 Jun 2024 |
somsagar07/FailureShiftRL/Baselines/Generation/config.py 2af53582222a76e3 |
ran
|
MIT (permissive) |
| LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study |
26 Apr 2024 |
aix-group/llms-for-cfs/src/gen_cf/generate_hatespeech.py 4bb808bb894ec755 |
ran
|
no licence file found · pointer only |
| LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study |
26 Apr 2024 |
aix-group/llms-for-cfs/src/gen_cf/generate_imdb.py 91e0ef1413933bfa |
ran
|
no licence file found · pointer only |
| LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study |
26 Apr 2024 |
aix-group/llms-for-cfs/src/gen_cf/generate_snli.py d94640c9148fbf50 |
unverified |
no licence file found · pointer only |
| Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant |
2024-03 (from id) |
dcase2024-task7-sound-scene-synthesis/fadtk/example/prompts/gpt4_quality.py ca4b05e19c89061a |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| RAGGED: Towards Informed Design of Retrieval Augmented Generation Systems |
14 Mar 2024 |
neulab/ragged/reader/utils.py 289d7efd299b90d4 |
ran
fingerprinted |
MIT (permissive) |
| An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4 |
5 Mar 2024 |
huihuichyan/unlimitedjudge/src/build_prompt_judge.py 5bb3a6f55ea7c8a7 |
ran
|
no licence file found · pointer only |
| Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation |
21 Feb 2024 |
pphuc25/distil-cd/src/dcd/prompts.py 8e4e00b93dabaf41 |
unverified |
no licence file found · pointer only |
| Visual Style Prompting with Swapping Self-Attention |
20 Feb 2024 |
naver-ai/Visual-Style-Prompting/vsp_real_script.py 56be14bdb80a0850 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Uncertainty Quantification for In-Context Learning of Large Language Models |
15 Feb 2024 |
lingchen0331/uq_icl/sources/utils.py 221aadea28e3360f |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Anchor-based Large Language Models |
12 Feb 2024 |
pangjh3/anllm/inference.py 4a62d89cb68ccb26 |
ran
|
no licence file found · pointer only |
| Anchor-based Large Language Models |
12 Feb 2024 |
pangjh3/anllm/analysis/get_infer_attention.py 65a7d55fbb2fbdf6 |
ran
|
no licence file found · pointer only |
| Anchor-based Large Language Models |
12 Feb 2024 |
pangjh3/anllm/applications/translation/inference_allmnoacinput.py 6ed5904743254ffa |
ran
|
no licence file found · pointer only |
| Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy |
11 Feb 2024 |
lmb-freiburg/ovqa/ovqa/create_prompt.py e1cf91cb1d73f4b7 |
ran
|
Apache-2.0 (permissive) |
| Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model |
23 Jan 2024 |
zwhe99/FeedbackMT/src/inference_sft.py 4a62d89cb68ccb26 |
ran
|
no licence file found · pointer only |
| Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models |
16 Jan 2024 |
pangjh3/llm4mt/train/inference.py 4a62d89cb68ccb26 |
ran
|
no licence file found · pointer only |
| Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models |
16 Jan 2024 |
pangjh3/llm4mt/train/attention_alignment_llama2.py 65a7d55fbb2fbdf6 |
ran
|
no licence file found · pointer only |
| Expert-guided Bayesian Optimisation for Human-in-the-loop Experimental Design of Known Systems |
5 Dec 2023 |
trsav/hitl-bo/bo/reccomender.py d16d3b2b4597fe7a |
ran
|
MIT (permissive) |
| Adapting Frechet Audio Distance for Generative Music Evaluation |
2 Nov 2023 |
identical code first harvested elsewhere ca4b05e19c89061a |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models |
31 Aug 2023 |
gersteinlab/biocoder/inference/final_batch_run.py d4614f1cd57cdba4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Improving Translation Faithfulness of Large Language Models via Augmenting Instructions |
24 Aug 2023 |
pppa2019/swie_overmiss_llm4mt/src/inference.py 258ba32f9e3eb1d9 |
ran
|
no licence file found · pointer only |
| Improving Translation Faithfulness of Large Language Models via Augmenting Instructions |
24 Aug 2023 |
pppa2019/swie_overmiss_llm4mt/utils/convert_alpaca_to_hf.py 54458e263ddc9222 |
ran
|
no licence file found · pointer only |
| Instruction Position Matters in Sequence Generation with Large Language Models |
23 Aug 2023 |
adaxry/post-instruction/test/inference.py 1ccc9517aa991b74 |
ran · our draft was wrong
|
no licence file found · pointer only |
| LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models |
15 Jul 2023 |
adianliusie/comparative-assessment/src/prompts/load_prompt.py a0709570914a889f |
ran · our draft was wrong
|
no licence file found · pointer only |
| Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations |
14 Nov 2022 |
swarnahub/explanationhardness/generate_gpt3_explanations.py 8b736157e0d6c782 |
unverified |
MIT (permissive) |
| Medical Image Understanding with Pretrained Vision Language Models: A Comprehensive Study |
30 Sep 2022 |
membrai/miu-vl/make_autopromptsv2.py 2b0f47de5fbaeb12 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Prompting as Probing: Using Language Models for Knowledge Base Construction |
23 Aug 2022 |
hemile/iswc-challenge/baseline.py 72ee5ba094d0c0bb |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Prompting as Probing: Using Language Models for Knowledge Base Construction |
23 Aug 2022 |
hemile/iswc-challenge/gpt3_baseline.py 6fd929182163484e |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2025.findings-acl.51 |
|
ModeEric/ORBIT-Llama/orbit/evaluation/astrobench_tests.py 8d33d2648796fc52 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.emnlp-industry.190 |
|
ismail31416/CAPSTONE/capstone/reasoning/prompts.py b1ef4f8c3585b0c5 |
unverified |
MIT (permissive) |
| arXiv:2024.findings-emnlp.365 |
|
UKPLab/m2qa/Experiments/LLM_evaluation/prompts.py 1f8df1be24e8f46c |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-acl.255 |
|
ShubhamKumarNigam/PredEx/Code/LLMs/Prompt_based_Inference/inference_LLAMA-2-7B_prediction.py 3b529eee45ab8e87 |
unverified |
MIT (permissive) |
| arXiv:2024.findings-acl.255 |
|
ShubhamKumarNigam/PredEx/Code/LLMs/Prompt_based_Inference/inference_LLAMA-2-7B_prediction_explanation.py 81749bb8b058f914 |
unverified |
MIT (permissive) |