| Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding added by Syntology |
2026-09 (from id) |
MinooAhmadii/chopthin-consensus-power-sampling/ccps/graders/math_normalize.py d51cc295307937e6 |
unverified |
licence not identified · pointer only |
| S$^2$Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models added by Syntology |
2026-09 (from id) |
yuanyuanjia71-spec/S2Prune/s2prune/vizwiz_eval.py 3bea898e987fab1e |
unverified |
no licence file found · pointer only |
| ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology |
2026-09 (from id) |
ZaiizaiZHANG/ISO-RAG/analyze_predictions.py 0d2dc85735d18403 |
unverified |
no licence file found · pointer only |
| ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation added by Syntology |
2026-09 (from id) |
ZaiizaiZHANG/ISO-RAG/analyze_retrieval_qa_transfer.py 1691191d75c4338c |
unverified |
no licence file found · pointer only |
| Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering added by Syntology |
2026-08 (from id) |
mm-air/AMR-Agent/evaluate.py 9fc3dafdabb9a23a |
ran
fingerprinted |
MIT (permissive) |
| EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval added by Syntology |
2026-08 (from id) |
XrazyMee/EviReform/src/evireform/metrics.py 77114e88a1e17b9b |
ran
fingerprinted |
Apache-2.0 (permissive) |
| The Embedder's Dilemma: LLMs Are Better, but at What Cost? added by Syntology |
2026-08 (from id) |
embeddings-benchmark/embedders-dilemma/llm_judge/evaluators/llm_rag_evaluator.py 97012cf96aa79266 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories added by Syntology |
2026-08 (from id) |
Changhao-Xiang/OpenVisTool/distill/filter/evaluate_answer_rule.py 24e4999eff337610 |
ran
fingerprinted |
MIT (permissive) |
| Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States added by Syntology |
2026-08 (from id) |
zazamrykh/internal_probing/src/dataset/squad_eval.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale added by Syntology |
2026-08 (from id) |
MBZUAI-CLeaR/IoE-Prompting/run_Hotpot_IoE.py b1a674018a8b1ccf |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale added by Syntology |
2026-08 (from id) |
MBZUAI-CLeaR/IoE-Prompting/run_math_IoE.py 5d6dae5b935c9916 |
ran
fingerprinted |
no licence file found · pointer only |
| When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO added by Syntology |
2026-08 (from id) |
CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO/domain_config.py 6151d15713a7392f |
ran
fingerprinted |
MIT (permissive) |
| CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents added by Syntology |
2026-08 (from id) |
lazzy1225/CoEvoKG/coevokg/reward/score/coevokg_score.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| ReFact: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning added by Syntology |
2026-07 (from id) |
NEUIR/REFACT/verl/verl/utils/reward_score/evdience_reward.py 445de90902805074 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation added by Syntology |
2026-07 (from id) |
tptrix29/RIMS/baseline/baseline_instructrag.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TrustMargin: Training-Free Arbitration between Parametric Memory and Retrieved Evidence in Large Language Models added by Syntology |
2026-06 (from id) |
mojixu/TrustMargin/src/trustmargin.py 3dc410d21f1c3741 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents added by Syntology |
2026-06 (from id) |
Ji-shuo/MRAgent/eval/evaluation.py e096279b3e0d7bd9 |
ran
fingerprinted |
no licence file found · pointer only |
| VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark added by Syntology |
2026-06 (from id) |
vampsbenchmark/VAMPS/extract_answers.py 5c64b982057e8460 |
ran
fingerprinted |
MIT (permissive) |
| Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling added by Syntology |
2026-06 (from id) |
kaist-cvml/perception-judge/prepare-datasets/utils.py 2b85173eac27fd50 |
ran
fingerprinted |
no licence file found · pointer only |
| OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources added by Syntology |
2026-05 (from id) |
JinheonBaek/OmniRetrieval/src/evaluation/metrics.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement added by Syntology |
2026-05 (from id) |
CuSO4-Chen/AKBE/AKBE/verl_akbe/verl/utils/reward_score/reward_em_betagrpo.py 6fcc2c2f4fe3ad67 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation added by Syntology |
2026-05 (from id) |
Yifan-Lan/zero-cot-probe/math_normalize.py 644c1663d2f9533f |
unverified |
MIT (permissive) |
| ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning added by Syntology |
2026-05 (from id) |
misonsky/chunk/metrics/cuad/compute_score.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MINTEVAL: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems added by Syntology |
2026-05 (from id) |
amy-hyunji/MINTEval/src/mem_alpha/memalpha/llm_agent/metrics.py 3b3fe5d04eaa0d33 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration added by Syntology |
2026-05 (from id) |
PKU-SEC-Lab/SPEX/evaluate/evaluate_utils/math_normalize.py 101b1e54fcdadf90 |
ran
fingerprinted |
no licence file found · pointer only |
| RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners added by Syntology |
2026-05 (from id) |
JugalGajjar/RSAT/src/rewards/answer_reward.py ee8c43ee40891b07 |
ran
fingerprinted |
MIT (permissive) |
| GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs added by Syntology |
2026-04 (from id) |
ulab-uiuc/GraphPlanner/router_planner/shared/utils.py c2e95e64eec380a8 |
ran
fingerprinted |
no licence file found · pointer only |
| Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall? added by Syntology |
2026-04 (from id) |
ultor1996/reasoning_primitives/src/utils.py 85ea1889d0093c0d |
ran
fingerprinted |
no licence file found · pointer only |
| OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models added by Syntology |
2026-04 (from id) |
LightChen233/OMIBench/omi_bench.py 9ea845c01a954867 |
ran · our draft was wrong
fingerprinted |
licence not identified · pointer only |
| Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning added by Syntology |
2026-04 (from id) |
koreankiwi99/formalization-gaming/src/utils/answer_parsing.py 651b0f7470e1a3e9 |
unverified |
MIT (permissive) |
| Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization added by Syntology |
2026-04 (from id) |
BITHLP/PolicySplit/verl/utils/reward_score/openai_math_grade.py 101b1e54fcdadf90 |
ran
fingerprinted |
no licence file found · pointer only |
| A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities added by Syntology |
2026-04 (from id) |
cjia7/DPR/src/npti/eval/eval_bbh.py caa05d10c796877f |
unverified |
no licence file found · pointer only |
| CASK: Core-Aware Selective KV Compression for Reasoning Traces added by Syntology |
2026-04 (from id) |
THUDM/LongBench/LongBench/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering added by Syntology |
2026-04 (from id) |
zhuyjan/WikiSeeker/utils/infoseek_evaluation_utils.py 4925192ddc8dfa9a |
unverified |
Apache-2.0 (permissive) |
| REAM: Merging Improves Pruning of Experts in LLMs added by Syntology |
2026-04 (from id) |
zai-org/glm-simple-evals/evals/grading/math_normalize.py 101b1e54fcdadf90 |
ran
fingerprinted |
no licence file found · pointer only |
| Crystal: Characterizing Relative Impact of Scholarly Publications added by Syntology |
2026-03 (from id) |
allenai/multicite/qa/qasper_evaluator.py d6eb7fc276a790c2 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models added by Syntology |
2026-03 (from id) |
XiaojieGu/Omanic/eval/open_eval.py 5cdd5f07b69b4933 |
unverified |
no licence file found · pointer only |
| Temporal Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN added by Syntology |
2026-03 (from id) |
terrierteam/temporal_conflicts/fact_mutability/analysis/f1_score.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Meta-Reinforcement Learning with Self-Reflection for Agentic Search added by Syntology |
2026-03 (from id) |
tengxiao1/MR-Search/verl/utils/reward_score/search_r1_like_qa_em.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Stacked from One: Multi-Scale Self-Injection for Context Window Extension added by Syntology |
2026-03 (from id) |
Clement25/SharedLLM/LongBenchTest/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Stacked from One: Multi-Scale Self-Injection for Context Window Extension added by Syntology |
2026-03 (from id) |
Clement25/SharedLLM/utils.py 6c668580324cfd97 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Spectral Attention Steering for Prompt Highlighting added by Syntology |
2026-03 (from id) |
waylonli/SEKA/pastalib/profiler/synthetic_profiler.py 76b695cd06f395b4 |
unverified |
Apache-2.0 (permissive) |
| SkillOrchestra: Learning to Route Agents via Skill Transfer added by Syntology |
2026-02 (from id) |
jiayuww/SkillOrchestra/skillorchestra/eval/metrics.py f23e31c1141ee0fe |
unverified |
Apache-2.0 (permissive) |
| Powering Up Zeroth-Order Training via Subspace Gradient Orthogonalization added by Syntology |
2026-02 (from id) |
OPTML-Group/ZO-Muon/llm/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Small Reward Models via Backward Inference added by Syntology |
2026-02 (from id) |
yikee/FLIP/metrics.py 5a9d2662f435748f |
unverified |
no licence file found · pointer only |
| CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering added by Syntology |
2026-02 (from id) |
How-Young-X/CompactRAG/src/metrics/F1Eval.py 26a63699755bd04e |
unverified |
MIT (permissive) |
| Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation added by Syntology |
2026-02 (from id) |
GasolSun36/BAR-RAG/filter.py 85fc7d4bd5e8711a |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Reinforcement Fine-Tuning for History-Aware Dense Retriever in RAG added by Syntology |
2026-02 (from id) |
zyc140345/HARR/metric/qa_em.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| GRAPHDANCER: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training added by Syntology |
2026-02 (from id) |
leopoldwhite/GraphDancer/verl/utils/reward_score/graphdancer_qa_em.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Alignment-Aware Model Adaptation via Feedback-Guided Optimization added by Syntology |
2026-02 (from id) |
facebookresearch/TruthRL/evaluation/evaluate.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity added by Syntology |
2026-02 (from id) |
AI21Labs/in-context-ralm/eval_qa.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| On the Paradoxical Interference between Instruction-Following and Task Solving added by Syntology |
2026-01 (from id) |
kijlk/IF-Interference/src/math_and_qa/qa_utils.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Textual Equilibrium Propagation for Deep Compound AI Systems added by Syntology |
2026-01 (from id) |
MinghuiChen43/TEP/src/tep/metrics/f1.py 03171efd2aab7e80 |
unverified |
Apache-2.0 (permissive) |
| SEEK: Steering LLM Reasoning for RAG via Internal Reasoning Sketches added by Syntology |
2026-01 (from id) |
OpenBMB/PAGER/src/evaluate_infer.py d22718d531542ab7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG added by Syntology |
2026-01 (from id) |
fengranMark/OpenDecoder/src/evaluation.py 24a31436a8dbd623 |
unverified |
MIT (permissive) |
| Revealing the Attention Floating Mechanism in Masked Diffusion Models added by Syntology |
2026-01 (from id) |
NEUIR/Attention-Floating/evaluate/evaluate_dream.py ff3b542378f5d19c |
unverified |
no licence file found · pointer only |
| HAPS: Hierarchical LLM Routing with Joint Architecture and Parameter Search added by Syntology |
2026-01 (from id) |
zihangtian/HAPS/HotpotQA/joint_rl/rl_batch.py 6c668580324cfd97 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval added by Syntology |
7 Jan 2026 |
ChenWangHKU/DTR/evaluation/metrics.py b92b07bb3574992a |
unverified |
no licence file found · pointer only |
| BaseCal: Unsupervised Confidence Calibration via Base Model Signals added by Syntology |
2026-01 (from id) |
Tan-Hexiang/BaseCal/gen_metric/em.py 9891ffc5ad4f4684 |
unverified |
no licence file found · pointer only |
| Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning added by Syntology |
2025-11 (from id) |
aiming-lab/Agent0/Agent0-VL/verl/evaluation/metrics.py e92c6c13f949ae88 |
unverified |
Apache-2.0 (permissive) |
| Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training added by Syntology |
2025-11 (from id) |
griver/Q-RAG/eval_llm_openqa.py b51d54e1d87d8bd9 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA added by Syntology |
2025-10 (from id) |
SapienzaNLP/LiteraryQA/literaryqa/ngram_metrics.py dfcf1201010b7d31 |
unverified |
no licence file found · pointer only |
| BOOTSTRAPPING LLMS TO REASON OVER LONGER HORIZONS VIA REINFORCEMENT LEARNING added by Syntology |
2025-10 (from id) |
AlesyaIvanova/h1/math_utils.py 101b1e54fcdadf90 |
ran
fingerprinted |
no licence file found · pointer only |
| ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards added by Syntology |
2025-10 (from id) |
identical code first harvested elsewhere dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Mamba Modulation On the Length Generalization of Mamba added by Syntology |
2025-09 (from id) |
gnepul-ace/mamba_modulation/Mamba/LongBench/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding added by Syntology |
2025-09 (from id) |
xiaobo-xing/TableDART/evaluate.py 90ce0a2b27b15e27 |
unverified |
MIT (permissive) |
| ConfTuner: Training Large Language Models to Express Their Confidence Verbally added by Syntology |
2025-08 (from id) |
liushiliushi/ConfTuner/src/llama_recipes/datasets2/hotpot_qa.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| arXiv:2507.21892 |
2025-07 (from id) |
LHRLAB/Graph-R1/evaluation/eval_r.py 78b0c6179667e101 |
unverified |
MIT (permissive) |
| The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models |
30 May 2025 |
nusnlp/fspo/evaluate/inference.py d22718d531542ab7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset |
27 May 2025 |
microsoft/rstar/fused_compute_score/prime_math/math_normalize.py 31113f93a6e0e98a |
unverified |
MIT (permissive) |
| Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers |
26 May 2025 |
mangopy/searchlm/src/utilize/metrics.py 93730b62e07a68f0 |
unverified |
Apache-2.0 (permissive) |
| R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning |
22 May 2025 |
RUCAIBox/R1-Searcher/evaluation/metric_calc_rule.py d22718d531542ab7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Fine-tuning Quantized Neural Networks with Zeroth-order Optimization |
19 May 2025 |
maifoundations/qzo/large_language_models/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Adaptive Orchestration of Modular Generative Information Access Systems |
24 Apr 2025 |
informagi/AQA/Adaptive-RAG/evaluate.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TTRL: Test-Time Reinforcement Learning |
22 Apr 2025 |
prime-rl/ttrl/verl/verl/utils/reward_score/ttrl_math/math_normalize.py 31113f93a6e0e98a |
unverified |
MIT (permissive) |
| Retrieval-Augmented Generation with Conflicting Evidence |
17 Apr 2025 |
hannight/ramdocs/run_madam_rag.py 42413e5aa5425a3e |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Out of Style: RAG's Fragility to Linguistic Variation |
11 Apr 2025 |
springcty/rag-fragility-to-linguistic-variation/LLM_generation/c_eval_generation.py bf618b8796ac4942 |
unverified |
MIT (permissive) |
| Single-Pass Document Scanning for Question Answering |
4 Apr 2025 |
mambaretriever/mambaretriever/rag_pipeline/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models? |
2 Apr 2025 |
bigai-ai/ToM-RL/verl/utils/reward_score/explore_tom.py 1436322b957b9b34 |
unverified |
MIT (permissive) |
| Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models? |
2 Apr 2025 |
bigai-ai/ToM-RL/verl/utils/reward_score/explore_tom_2.py c5322e5a65a9bea1 |
unverified |
MIT (permissive) |
| Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models |
25 Mar 2025 |
stogiannidis/srbench/src/eval/acc.py 117783c12219bab5 |
unverified |
MIT (permissive) |
| Radar: Fast Long-Context Decoding for Any Transformer |
13 Mar 2025 |
borealisai/radar-decoding/research/longbench/longbench.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
LGPL-3.0 (copyleft) · pointer only |
| Agent models: Internalizing Chain-of-Action Generation into Reasoning models |
9 Mar 2025 |
adam-bjtu/autocoa/evaluation/metrics.py 98a43410c5681732 |
unverified |
MIT (permissive) |
| Words or Vision: Do Vision-Language Models Have Blind Faith in Text? |
4 Mar 2025 |
d-ailin/blind-faith-in-text/raw_code/hf_evaluator.py 0bcff406512311b4 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling |
10 Feb 2025 |
RyanLiu112/compute-optimal-tts/src/envs/MATH/verify_utils.py 101b1e54fcdadf90 |
ran
fingerprinted |
MIT (permissive) |
| C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation |
10 Feb 2025 |
Chen-GX/C-3PO/C-3PO/metrics.py 37e0a25b59a4cc3e |
unverified |
Apache-2.0 (permissive) |
| Process Reinforcement through Implicit Rewards |
3 Feb 2025 |
prime-rl/prime/data_preprocessing/math_util/math_normalize.py 101b1e54fcdadf90 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home |
22 Jan 2025 |
s-nlp/AdaRAGUE/Adaptive_Rag/evaluate.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Tell me about yourself: LLMs are aware of their learned behaviors |
19 Jan 2025 |
xuchanbao/behavioral-self-awareness/code/multiple-choice/evaluation/self-report-free-separate-rephrasings.py b75e3691451b3c33 |
unverified |
MIT (permissive) |
| Context-DPO: Aligning Language Models for Context-Faithfulness |
18 Dec 2024 |
identical code first harvested elsewhere e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Foundational Large Language Models for Materials Research |
12 Dec 2024 |
M3RG-IITD/llamat/src/Kshot-val.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| SiReRAG: Indexing Similar and Related Information for Multihop Reasoning |
9 Dec 2024 |
SalesforceAIResearch/SiReRAG/evaluate_2wiki.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Free Process Rewards without Process Labels |
2 Dec 2024 |
lifan-yuan/implicitprm/eval/math_utils/math_normalize.py 101b1e54fcdadf90 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning |
25 Nov 2024 |
THU-KEG/AtomR/src/calculate_metrics.py 702b1281050dbc40 |
unverified |
no licence file found · pointer only |
| Squeezed Attention: Accelerating Long Context Length LLM Inference |
14 Nov 2024 |
SqueezeAILab/SqueezedAttention/LongBench/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Large Language Models Can Self-Improve in Long-context Reasoning |
12 Nov 2024 |
sihengli99/sealong/eval_longbench_qa.py dae7ab386661a4f4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning |
25 Oct 2024 |
fyyfu/headkv/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering |
23 Oct 2024 |
QingFei1/LongRAG/src/metric.py c6a80c065d2e4851 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning |
22 Oct 2024 |
zmzhang2000/trustworthy-alignment/dschat/utils/metrics.py 9706a9a0b961096b |
unverified |
no licence file found · pointer only |
| SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction |
17 Oct 2024 |
sail-sg/simlayerkv/LongBench/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models |
17 Oct 2024 |
hoar012/rap-mllm/eval/eval_qa.py 7927b07dcc328a1f |
ran
fingerprinted |
no licence file found · pointer only |
| Synthetic Knowledge Ingestion: Towards Knowledge Refinement and Injection for Enhancing Large Language Models |
12 Oct 2024 |
intuit-ai-research/knowledge-infused-ai/synthetic-knowledge-ingestion/finetuning/ft_llama_factory.py 6a96435eba311b08 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Zeroth-Order Fine-Tuning of LLMs in Random Subspaces |
11 Oct 2024 |
zimingyy/subzero/large_models/metrics.py e7e75981cb464788 |
ran · our draft was wrong
fingerprinted |
GPL-3.0 (copyleft) · pointer only |
| StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models |
10 Oct 2024 |
kmc0207/Stableprompt/utils.py d8af265929dc927f |
ran
fingerprinted |
no licence file found · pointer only |