| Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning added by Syntology |
2026-09 (from id) |
TonyQJH/cca-emnlp2026/code/cca_core.py 12de2cde1bb6e417 |
unverified |
Apache-2.0 (permissive) |
| CogEvol: Towards Efficient and Reliable Learning Environment Generation added by Syntology |
2026-08 (from id) |
CogEvol/CogEvol-4B/eval/slide_eval.py b2d30690ed8eef53 |
unverified |
MIT (permissive) |
| Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery added by Syntology |
2026-07 (from id) |
SKYLENAGE-AI/HypoArena/basics/parsing.py 2f41dc825738de71 |
ran
|
no licence file found · pointer only |
| Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? added by Syntology |
2026-06 (from id) |
THU-KEG/RuVerBench/code/run_judges/agenticcoding_runner.py 46c0327ba1728a45 |
ran
|
licence not identified · pointer only |
| Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? added by Syntology |
2026-06 (from id) |
THU-KEG/RuVerBench/code/run_judges/deepresearch_runner.py 38372855661357c6 |
ran
|
licence not identified · pointer only |
| Argument Collapse: LLMs Flatten Long-Form Public Debate added by Syntology |
2026-06 (from id) |
mungg/argument_collapse/src/argument_collapse/annotate/pair_comparison_main_arg.py b933cab935ecc2aa |
ran
|
MIT (permissive) |
| Argument Collapse: LLMs Flatten Long-Form Public Debate added by Syntology |
2026-06 (from id) |
mungg/argument_collapse/src/argument_collapse/annotate/stance.py ec41fe411ec4206b |
ran
|
MIT (permissive) |
| Argument Collapse: LLMs Flatten Long-Form Public Debate added by Syntology |
2026-06 (from id) |
mungg/argument_collapse/src/argument_collapse/annotate/structure.py ba2ef2aa84eacd6c |
ran
|
MIT (permissive) |
| From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models added by Syntology |
2026-04 (from id) |
microsoft/PairedSafety/analysis/grader_error_analysis/llm_code_errors.py f430bb03b2488d92 |
ran
|
MIT (permissive) |
| Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall? added by Syntology |
2026-04 (from id) |
ultor1996/reasoning_primitives/src/utils.py e4135e8bf295baaf |
ran
fingerprinted |
no licence file found · pointer only |
| Do Phone-Use Agents Respect Your Privacy? added by Syntology |
2026-04 (from id) |
FreedomIntelligence/MyPhoneBench/android_world/agents/agent_utils.py 382a844d97026811 |
unverified |
Apache-2.0 (permissive) |
| Memento-Skills: Let Agents Design Agents Memento-Team added by Syntology |
2026-03 (from id) |
Memento-Teams/Memento-Skills/core/memento_s/utils.py 4b48b517a30002c3 |
unverified |
Apache-2.0 (permissive) |
| RoboLayout: Differentiable 3D Scene Generation for Embodied Agents added by Syntology |
2026-03 (from id) |
alishams21/robolayout/src/layoutvlm/layoutvlm.py b0987cb5d12e53ca |
unverified |
no licence file found · pointer only |
| According to Me: Long-Term Personalized Referential Memory QA added by Syntology |
2026-03 (from id) |
JingbiaoMei/ATM-Bench/memqa/qa_agent_baselines/A-Mem/memory_layer.py 1ba9020de9ace83e |
unverified |
MIT (permissive) |
| MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering added by Syntology |
2026-02 (from id) |
AIDASLab/MATA/utils/JA_extract_answer.py 36d4245c8e43ef52 |
unverified |
no licence file found · pointer only |
| CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering added by Syntology |
2026-02 (from id) |
How-Young-X/CompactRAG/src/core/AskCorpus.py 9f0a5786975dd602 |
unverified |
MIT (permissive) |
| CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval added by Syntology |
2026-01 (from id) |
idirlab/CaseFacts/experiments/check_claims_contradiction.py 6278226e597fa860 |
unverified |
no licence file found · pointer only |
| CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval added by Syntology |
2026-01 (from id) |
idirlab/CaseFacts/experiments/check_diff_case_contradiction.py 678ae97e3c9c6232 |
unverified |
no licence file found · pointer only |
| QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems added by Syntology |
2025-12 (from id) |
yyiliu/QuadSentinel/src/quadsentinel/utils/functions.py 402263fe04b4fc5c |
unverified |
no licence file found · pointer only |
| OptiTree: Hierarchical Thoughts Generation with Tree Search for LLM Optimization Modeling added by Syntology |
2025-10 (from id) |
MIRALab-USTC/OptiTree/utils.py 3db55b75deebf780 |
unverified |
MIT (permissive) |
| MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations |
25 Jun 2025 |
mirage-benchmark/mirage-benchmark/MMMT/src/clarification_generator.py 9593836b0ce2c31a |
unverified |
no licence file found · pointer only |
| ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback |
23 May 2025 |
EnVision-Research/ComfyMind/utils/tools.py 00e58640391e7193 |
unverified |
MIT (permissive) |
| LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics |
30 Apr 2025 |
marc1198/chat-hsr/without-ROS/main_single_agent.py 02d7798ffd79f86d |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework |
30 Apr 2025 |
Miracle1207/Mean-Field-LLM/mf_llm/evaluate/evaluate_gpt_batch.py 856165cff8587e67 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents |
25 Feb 2025 |
Alibaba-NLP/ViDoRAG/vidorag_agents.py 34692b533654cd8b |
ran · our draft was wrong
|
no licence file found · pointer only |
| Atom of Thoughts for Markov LLM Test-Time Scaling |
17 Feb 2025 |
qixucen/atom/experiment/module.py 41aebb57ad1d527d |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents |
17 Oct 2024 |
damo-nlp-sg/coi-agent/utils.py 8da3d9ba5f65d0b8 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation |
7 Oct 2024 |
opengvlab/phygenbench/PhyGenEval/semantic/gpt4o_sementic.py f7510953d3aad7af |
ran
fingerprinted |
no licence file found · pointer only |
| T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation |
19 Jul 2024 |
KaiyueSun98/T2V-CompBench/LLaVA/llava/eval/compbench_eval_action_binding.py 18e1081b616ae0e0 |
ran
fingerprinted |
no licence file found · pointer only |
| T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation |
19 Jul 2024 |
KaiyueSun98/T2V-CompBench/LLaVA/llava/eval/compbench_eval_consistent_attr.py 0b9124dcf2830cbb |
ran
|
no licence file found · pointer only |
| ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents |
28 Jun 2024 |
EachSheep/ShortcutsBench/experiments/all_experiments.py adb0f88a9d1082c2 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents |
23 May 2024 |
google-research/android_world/android_world/agents/agent_utils.py 382a844d97026811 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM |
9 May 2024 |
yancykahn/coa/common.py 9f289a28c6cc9b87 |
ran
fingerprinted |
MIT (permissive) |
| Uncovering Safety Risks of Large Language Models through Concept Activation Vector |
18 Apr 2024 |
tmlr-group/DeepInception/common.py d7a1e728e5b74d3f |
ran
fingerprinted |
MIT (permissive) |
| Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks |
2 Apr 2024 |
tml-epfl/llm-adaptive-attacks/common.py d7a1e728e5b74d3f |
ran
fingerprinted |
MIT (permissive) |
| Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes |
1 Mar 2024 |
RICommunity/TAP/common.py 7337ffebcca7ec8c |
ran
fingerprinted |
MIT (permissive) |
| Defending LLMs against Jailbreaking Attacks via Backtranslation |
26 Feb 2024 |
yihanwang617/llm-jailbreaking-defense-backtranslation/PAIR/common.py d7a1e728e5b74d3f |
ran
fingerprinted |
BSD-3-Clause (permissive) |
| DeepInception: Hypnotize Large Language Model to Be Jailbreaker |
6 Nov 2023 |
tmlr-group/deepinception/common.py d7a1e728e5b74d3f |
ran
fingerprinted |
MIT (permissive) |
| Jailbreaking Black Box Large Language Models in Twenty Queries |
12 Oct 2023 |
patrickrchao/jailbreakingllms/conversers.py 7f22816d570dc592 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| GameEval: Evaluating LLMs on Conversational Games |
19 Aug 2023 |
gameeval/gameeval/chat/text003_chat.py e8aa7c0e2fd87cd4 |
ran
fingerprinted |
no licence file found · pointer only |