parse_score
parse_score appears in the code Syntology harvested for 49 papers, as 14 distinct code bodies found in 51 places (a place is one code body under one paper). At least one of them ran in 48 of the papers; 10 of the code bodies carry a behaviour fingerprint.
What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named parse_score do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.
Samples Syntology
Syntology ran 13 of the 14 distinct code bodies named parse_score; 1 is unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:
Licence is a property of each copy, so it is counted per place: 30 of the 51 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.
“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.
Papers
49 papers shown of 49, newest first; 51 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 2 papers added by Syntology; 1 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
| Paper | Date | File | Status Syntology | Licence |
|---|---|---|---|---|
| DELULU: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks added by Syntology | 2026-05 (from id) | microsoft/delulu/evaluations/run_delulu_judges.py 85a1174c53ba7c3a |
ran fingerprinted | MIT (permissive) |
| Psychological Steering of Large Language Models added by Syntology | 2026-04 (from id) | kaistAI/FLASK/gpt_review/gpt4_eval.py 933335551c0e8a52 |
ran · our draft was wrong fingerprinted | no licence file found · pointer only |
| LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning | 19 Mar 2025 | aimagelab/LLaVA-MORE/src/llava/eval/eval_gpt_review.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Re-Imagining Multimodal Instruction Tuning: A Representation View | 2 Mar 2025 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases | 3 Dec 2024 | kki2eve/agri-llava/agri_llava/eval/eval_gpt_review_visual.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| $\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models | 23 Nov 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Right this way: Can VLMs Guide Us to See More to Answer Questions? | 1 Nov 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Improve Vision Language Model Chain-of-thought Reasoning | 21 Oct 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge | 3 Oct 2024 | amazon-science/beyondcorrelation/examples/judge_bench_example.py 0b9867bdf14b90fb |
ran · our draft was wrong | Apache-2.0 (permissive) |
| FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models | 3 Oct 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Law of the Weakest Link: Cross Capabilities of Large Language Models | 30 Sep 2024 | facebookresearch/llm-cross-capabilities/evaluation/evaluate_response.py f73ce4685c268b2d |
ran fingerprinted | licence not identified · pointer only |
| SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding | 28 Aug 2024 | dptech-corp/Uni-SMART/SciLitLLM/cpt/quality_control/llama_infer.py 978a0952e0e705b2 |
ran fingerprinted | MIT (permissive) |
| ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models | 25 Aug 2024 | yejipark-m/convis/eval/eval_gpt_review_bench.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | MIT (permissive) |
| Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs | 2 Aug 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model | 23 Jul 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs | 17 Jun 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning | 17 Jun 2024 | naver-ai/elva/Elva/eval_gpt_review_parsing_bench.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | no licence file found · pointer only |
| Phased Instruction Fine-Tuning for Large Language Models | 1 Jun 2024 | xubuvd/phasedsft/evaluation/win_tie_loss_stat.py d9b9b1f09f63f865 |
ran fingerprinted | Apache-2.0 (permissive) |
| Descriptive Image Quality Assessment in the Wild | 29 May 2024 | XPixelGroup/DepictQA/src/eval/cal_gpt4_score_detail_v1.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Descriptive Image Quality Assessment in the Wild | 29 May 2024 | XPixelGroup/DepictQA/src/eval/cal_gpt4_score_detail_v2.py d22ddef4b1ae67fd |
ran fingerprinted | Apache-2.0 (permissive) |
| RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness | 27 May 2024 | rlhf-v/rlhf-v/eval/eval_gpt_review_llava.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | no licence file found · pointer only |
| LOVA3: Learning to Visual Question Answering, Asking and Assessment | 23 May 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| MoVA: Adapting Mixture of Vision Experts to Multimodal Context | 19 Apr 2024 | TempleX98/MoVA/mova/eval/eval_gpt_review.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models | 19 Apr 2024 | FoundationVision/Groma/groma/eval/eval_gpt_review_visual.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Beyond Embeddings: The Promise of Visual Table in Visual Reasoning | 27 Mar 2024 | lavi-lab/visual-table/llava/eval/eval_gpt_review_visual.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning | 25 Mar 2024 | deepcs233/visual-cot/llava/eval/eval_gpt_review_visual.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models | 14 Mar 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment | 21 Feb 2024 | hitsz-tmg/cognitive-visual-language-mapper/LLaVA/llava/eval/eval_gpt_review_visual.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | no licence file found · pointer only |
| Event-level Knowledge Editing | 20 Feb 2024 | thu-keg/event-level-knowledge-editing/examiner/parse_result.py 9e0abb69457fe1bc |
ran | licence not identified · pointer only |
| Event-level Knowledge Editing | 20 Feb 2024 | thu-keg/event-level-knowledge-editing/examiner/parse_result_locality.py 7da1f09a1752b039 |
ran | licence not identified · pointer only |
| Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models | 16 Feb 2024 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| ODIN: Disentangled Reward Mitigates Hacking in RLHF | 11 Feb 2024 | identical code first harvested elsewhere 8049b382893c73dd |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning | 7 Feb 2024 | identical code first harvested elsewhere 8049b382893c73dd |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Osprey: Pixel Understanding with Visual Instruction Tuning | 15 Dec 2023 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM | 14 Dec 2023 | identical code first harvested elsewhere 8049b382893c73dd |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Hallucination Augmented Contrastive Learning for Multimodal Large Language Model | 12 Dec 2023 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos | 7 Dec 2023 | aldraus/quilt-llava/llava/eval/quilt_gpt_eval.py 763f6fe20bd52fbb |
ran · violated contract fingerprinted | MIT (permissive) |
| Alpha-CLIP: A CLIP Model Focusing on Wherever You Want | 6 Dec 2023 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| The Philosopher's Stone: Trojaning Plugins of Large Language Models | 2023-12 (from id) | chichidd/llm-lora-trojan/eval/gpt_review.py 8049b382893c73dd |
ran · violated contract fingerprinted | Apache-2.0 (permissive) |
| Improved Baselines with Visual Instruction Tuning | 5 Oct 2023 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models | 18 Sep 2023 | identical code first harvested elsewhere 763f6fe20bd52fbb |
ran · violated contract fingerprinted | licence of this copy not recorded |
| UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition | 7 Aug 2023 | universal-ner/universal-ner/src/train/fastchat/eval/eval_gpt_review.py 8049b382893c73dd |
ran · violated contract fingerprinted | MIT (permissive) |
| FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets | 20 Jul 2023 | kaistai/flask/gpt_review/gpt4_eval.py 933335551c0e8a52 |
ran · our draft was wrong fingerprinted | no licence file found · pointer only |
| AlpaGasus: Training A Better Alpaca with Fewer Data | 17 Jul 2023 | gpt4life/alpagasus/rating/filter.py b748891b4b063b83 |
ran · violated contract fingerprinted | no licence file found · pointer only |
| PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations | 6 Jul 2023 | bcdnlp/prd/peer_rank/eval_bard_review.py 8049b382893c73dd |
ran · violated contract fingerprinted | MIT (permissive) |
| QLoRA: Efficient Finetuning of Quantized LLMs | 23 May 2023 | artidoro/qlora/eval/eval_gpt_review.py 8049b382893c73dd |
ran · violated contract fingerprinted | MIT (permissive) |
| Lion: Adversarial Distillation of Proprietary Large Language Models | 22 May 2023 | yjiangcm/lion/src/chatgpt_referee.py eb6629df9a940cb4 |
unverified | MIT (permissive) |
| Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding | 19 May 2023 | bowang-lab/clinical-camel/evaluation/eval_gpt_review.py 8049b382893c73dd |
ran · violated contract fingerprinted | AGPL-3.0 (copyleft) · pointer only |
| G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment | 29 Mar 2023 | megagonlabs/llm-longeval/eval/evaluator.py f662bdaf543ab06e |
ran · honoured contract fingerprinted | BSD-3-Clause (permissive) |
| GPT-4 Technical Report | 15 Mar 2023 | identical code first harvested elsewhere b748891b4b063b83 |
ran · violated contract fingerprinted | licence of this copy not recorded |
| arXiv:2024.findings-acl.341 | xubuvd/PhasedSFT/evaluation/win_tie_loss_stat.py d9b9b1f09f63f865 |
ran fingerprinted | Apache-2.0 (permissive) |
This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections