| Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks added by Syntology |
2026-09 (from id) |
Jaong/CreaEval/experiment.py 157153a5f87f4a2d |
unverified |
no licence file found · pointer only |
| Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs added by Syntology |
2026-09 (from id) |
yuntian-group/cdsep/cdsep/schema.py 3e955ba82b75f281 |
unverified |
MIT (permissive) |
| Auditing Alignment Controllability in LLMs via Political Axes added by Syntology |
2026-07 (from id) |
mbrcic/llm-political-steerability/ingestion/01_collect/collect.py bcbcfcee1dc27141 |
ran
fingerprinted |
MIT (permissive) |
| Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian added by Syntology |
2026-06 (from id) |
DS4AI-UPB/crosslingual-romanian-re/infer_e2e.py ffe27e79dc4e4979 |
ran
fingerprinted |
no licence file found · pointer only |
| RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning added by Syntology |
2026-06 (from id) |
jwang-580/RADIANT-PET/eval/infer_api_models.py 88548e534a655e64 |
ran
fingerprinted |
no licence file found · pointer only |
| RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning added by Syntology |
2026-06 (from id) |
jwang-580/RADIANT-PET/eval/infer_gpt_oss.py 6a5a10fbcab9e3a1 |
ran
fingerprinted |
no licence file found · pointer only |
| RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning added by Syntology |
2026-06 (from id) |
jwang-580/RADIANT-PET/eval/infer_medgemma.py 0c270d00c05c496c |
ran
fingerprinted |
no licence file found · pointer only |
| SMCEVOLVE: Principled Scientific Discovery via Sequential Monte Carlo Evolution added by Syntology |
2026-05 (from id) |
kongwanbianjinyu/SMCEvolve/smcevolve/prompts.py 47e92f4ea3de4ddf |
ran
|
no licence file found · pointer only |
| Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models added by Syntology |
2026-05 (from id) |
leeyejin1231/DLM_Steering_Remasking/utils/mmlu_eval.py bae3b10ba5297700 |
ran
fingerprinted |
no licence file found · pointer only |
| Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen added by Syntology |
2026-04 (from id) |
synthiumjp/koriat/collect_data.py 9cf19b493bfa49a0 |
ran
fingerprinted |
MIT (permissive) |
| Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen added by Syntology |
2026-04 (from id) |
synthiumjp/koriat/collect_data_v2.py 51e6f27b5ffb8628 |
ran
fingerprinted |
MIT (permissive) |
| THIVLVC: Retrieval Augmented Dependency Parsing for Latin added by Syntology |
2026-04 (from id) |
l-pommeret/THIVLVC/src/pipeline/stage1_rag_with_baseline.py a7e44ce956d21681 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Alignment-Aware Model Adaptation via Feedback-Guided Optimization added by Syntology |
2026-02 (from id) |
facebookresearch/TruthRL/evaluation/evaluate.py 21fa52735b227537 |
unverified |
licence not identified · pointer only |
| Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation added by Syntology |
2026-01 (from id) |
sunxin000/TTARAG/local_evaluation.py d5f4dbf6439299f0 |
unverified |
no licence file found · pointer only |
| JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation added by Syntology |
2026-01 (from id) |
lhl/liquid-ai-hackathon-tokyo/eval/judge-mt.py 5b3a1e4d0aa325d5 |
unverified |
no licence file found · pointer only |
| AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite added by Syntology |
2025-10 (from id) |
allenai/agent-baselines/agent_baselines/solvers/code_agent/llm_agent.py cdad0f257db7572a |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.04127 |
2025-07 (from id) |
awslabs/graphrag-toolkit/byokg-rag/src/graphrag_toolkit/byokg_rag/byokg_query_engine.py 0cac8c1bedf394c9 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning |
2025-04 (from id) |
DataArcTech/SQL-R1/src/inference.py 4127f7104f276d1b |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs |
27 Mar 2025 |
CVI-SZU/FaceBench/evaluation/inference.py a4d72c2b03733d9b |
unverified |
MIT (permissive) |
| BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology |
28 Feb 2025 |
Future-House/BixBench/bixbench/utils.py a6c564d5b5c99437 |
unverified |
Apache-2.0 (permissive) |
| EgoNormia: Benchmarking Physical Social Norm Understanding |
27 Feb 2025 |
open-social-world/egonormia/src/gen/03_gen_questions.py a78ab5f88a9698d5 |
unverified |
MIT (permissive) |
| VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model |
26 Feb 2025 |
microsoft/gui-agent-rl/utils.py ff74d412d9163765 |
unverified |
MIT (permissive) |
| No Preference Left Behind: Group Distributional Preference Optimization |
28 Dec 2024 |
BigBinnie/GDPO/evaluate_BPC.py 80adbd3522975425 |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| Find Any Part in 3D |
20 Nov 2024 |
ziqi-ma/find3d/dataengine/llm/name_single_part_gemini.py d213559236b7abdd |
unverified |
MIT (permissive) |
| Find Any Part in 3D |
20 Nov 2024 |
ziqi-ma/find3d/dataengine/llm/query_orientation.py 1de3152e134b486b |
unverified |
MIT (permissive) |
| Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad? |
25 Oct 2024 |
ml-research/bongard-in-wonderland/experiments/evaluate/eval_bp_with_solutions.py 3580d02883c8fc63 |
unverified |
no licence file found · pointer only |
| Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning |
21 Oct 2024 |
kuan2jiu99/audio-hallucination/icassp2025/evaluation.py 698b1eaa3d1639e4 |
unverified |
no licence file found · pointer only |
| Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning |
21 Oct 2024 |
kuan2jiu99/audio-hallucination/interspeech2024/evaluation.py 94016dbb5d97a3d5 |
unverified |
no licence file found · pointer only |
| SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories |
11 Sep 2024 |
allenai/super-benchmark/super/agent/agent.py 7365c17255307e51 |
ran
|
Apache-2.0 (permissive) |
| Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos |
26 Aug 2024 |
qirui-chen/MultiHop-EgoQA/benchmark/metrics/evaluate_answering.py 19864996c48943c7 |
ran · our draft was wrong
|
no licence file found · pointer only |
| View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior |
1 Jul 2024 |
Bluefin-Tuna/ApartResearch/deception/pyfiles/agent.py 7d64ba24b41b4639 |
ran
|
MIT (permissive) |
| CRAG -- Comprehensive RAG Benchmark |
7 Jun 2024 |
facebookresearch/CRAG/local_evaluation.py a7323b55a4495886 |
ran · our draft was wrong
fingerprinted |
licence not identified · pointer only |
| SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales |
31 May 2024 |
xu1868/SaySelf/utils/utils.py df9fb4e8e291c1f4 |
ran
fingerprinted |
MIT (permissive) |
| Fleet of Agents: Coordinated Problem Solving with Large Language Models |
7 May 2024 |
au-clan/FoA/src/agents/crosswords.py da3bdfa23d2582f6 |
ran
fingerprinted |
MIT (permissive) |
| Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models |
17 Apr 2024 |
cicl-stanford/moral-evals/offtherails/src/evaluate_llm.py 0e5309d0f2137cd8 |
ran
fingerprinted |
MIT (permissive) |
| FABLES: Evaluating faithfulness and content selection in book-length summarization |
1 Apr 2024 |
lilakk/booookscore/booookscore/legacy/get_booookscore_v2.py 710b33af888f5924 |
ran
|
MIT (permissive) |
| DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation |
4 Mar 2024 |
shirley-wu/daco/evaluation/eval_helpfulness.py 9bd2a3e45b17af18 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually |
19 Jan 2024 |
SecureAIAutonomyLab/ConditionalVLM/data_scripts/step2_human_verification.py d7840e6c2b484338 |
ran
|
MIT (permissive) |
| MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI |
27 Nov 2023 |
eric-ai-lab/probmed/eval/calculate_score.py eabb48a101284727 |
unverified |
MIT (permissive) |
| GENOME: GenerativE Neuro-symbOlic visual reasoning by growing and reusing ModulEs |
8 Nov 2023 |
umass-foundation-model/genome/engine/gpt.py c949723a39704cc2 |
unverified |
Apache-2.0 (permissive) |
| Moral Foundations of Large Language Models |
23 Oct 2023 |
abdulhaim/moral_foundations_llm/utils/gpt3_utils.py 0f6ba13b6fc5cda6 |
ran
|
no licence file found · pointer only |
| BooookScore: A systematic exploration of book-length summarization in the era of LLMs |
1 Oct 2023 |
lilakk/BooookScore/booookscore/legacy/get_booookscore_v2.py 710b33af888f5924 |
ran
|
MIT (permissive) |
| Red Teaming Language Model Detectors with Language Models |
31 May 2023 |
shizhouxing/Attack-LM-Detectors/DetectGPT/openai_perturbations.py 0c125fa85591bb34 |
unverified |
BSD-3-Clause (permissive) |
| Pretraining Language Models with Human Preferences |
16 Feb 2023 |
tomekkorbak/pretraining-with-human-feedback/red_team.py 8d9e4494a1a48703 |
unverified |
MIT (permissive) |