| Fine-Tuning of Transformer models with Frames added by Syntology |
2026-08 (from id) |
vsingh-group/FrameFT/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference added by Syntology |
2026-08 (from id) |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| MEMORYCARD: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering added by Syntology |
2026-06 (from id) |
NEUIR/MemoryCard/lmms_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning added by Syntology |
2026-02 (from id) |
LivingFutureLab/CaCoVID/lmms_eval/lmms_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ρ-EOS: Training-free Bidirectional Variable-Length Control for Masked Diffusion LLMs added by Syntology |
2026-01 (from id) |
yjyddq/rho-EOS/dllm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs added by Syntology |
2026-01 (from id) |
wzj1718/PlaM/Vision-Token-Masking/lmms_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| CD 4 LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models added by Syntology |
2026-01 (from id) |
yihao-liang/CDLM/evaluation/dllm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior added by Syntology |
2025-12 (from id) |
yu-lin-li/DyToK/eval/lmms_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition |
19 May 2025 |
identical code first harvested elsewhere ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora |
10 Apr 2025 |
babylm/evaluation-pipeline/lm_eval/api/model.py de4e47a5e5126171 |
unverified |
MIT (permissive) |
| VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning |
9 Apr 2025 |
opengvlab/videochat-r1/Videochat-R1/lmms-eval_videochat/lmms_eval/api/model.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs |
7 Apr 2025 |
sunblaze-ucb/llm-api-audit/benchmark/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable |
1 Mar 2025 |
git-disl/safety-tax/eval/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data |
19 Feb 2025 |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Small Models Struggle to Learn from Strong Reasoners |
17 Feb 2025 |
Small-Model-Gap/Small-Model-Learnability-Gap/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Large Language Diffusion Models |
14 Feb 2025 |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| s1: Simple test-time scaling |
31 Jan 2025 |
simplescaling/s1/eval/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces |
18 Dec 2024 |
vision-x-nyu/thinking-in-space/lmms_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment |
15 Nov 2024 |
GATECH-EIC/AmoebaLLM/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| On the Loss of Context-awareness in General Instruction Fine-tuning |
5 Nov 2024 |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild |
7 Oct 2024 |
Model-GLUE/Model-GLUE/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding |
22 Sep 2024 |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| LIME: Less Is More for MLLM Evaluation |
10 Sep 2024 |
kangreen0210/lime/lmms_eval/api/model.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios |
30 Aug 2024 |
opendatalab/urbench/lmms_eval/api/model.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| On Large Language Model Continual Unlearning |
14 Jul 2024 |
identical code first harvested elsewhere 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation |
4 Jul 2024 |
mansicer/Q-Adapter/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models |
24 Jun 2024 |
abdelfattah-lab/shadow_llm/llm-interpret/lm-evaluation-harness/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Tokenization Falling Short: On Subword Robustness in Large Language Models |
17 Jun 2024 |
floatai/tkeval/evaluation/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| BlockPruner: Fine-grained Pruning for Large Language Models |
15 Jun 2024 |
MrGGLS/BlockPruner/lm_eval/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| BlockPruner: Fine-grained Pruning for Large Language Models |
15 Jun 2024 |
MrGGLS/BlockPruner/lm_eval/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| A safety realignment framework via subspace-oriented model fusion for large language models |
15 May 2024 |
xinykou/safety_realignment/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| EconLogicQA: A Question-Answering Benchmark for Evaluating Large Language Models in Economic Sequential Reasoning |
13 May 2024 |
yinzhu-quan/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation |
1 Apr 2024 |
hdong920/griffin/src/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| A Comprehensive Evaluation of Quantization Strategies for Large Language Models |
26 Feb 2024 |
cordercorder/quant_eval/quant/SpQR/lm-evaluation-harness/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| OneBit: Towards Extremely Low-bit Large Language Models |
17 Feb 2024 |
xuyuzhuang11/OneBit/evaluation/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference |
14 Feb 2024 |
hdong920/less/src/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security |
26 Dec 2023 |
zefang-liu/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| LAiW: A Chinese Legal Large Language Models Benchmark |
9 Oct 2023 |
dai-shen/laiw/src/financial-evaluation/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Do prompt positions really matter? |
23 May 2023 |
milliemaoo/prompt-position/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| RPTQ: Reorder-based Post-training Quantization for Large Language Models |
3 Apr 2023 |
hahnyuan/rptq4llm/models/models_utils.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams |
29 Mar 2023 |
piresramon/gpt-4-enem/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| A Few More Examples May Be Worth Billions of Parameters |
8 Oct 2021 |
yuvalkirstain/lm-evaluation-harness/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Language Models are Few-Shot Learners |
28 May 2020 |
Sypherd/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Language Models are Few-Shot Learners |
28 May 2020 |
vilm-ai/viet-llm-eval/lm_eval/api/model.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Language Models are Few-Shot Learners |
28 May 2020 |
neuralmagic/lm-evaluation-harness/lm_eval/api/model.py 26519d22d8af6d48 |
unverified |
MIT (permissive) |
| Language Models are Few-Shot Learners |
28 May 2020 |
EleutherAI/lm_evaluation_harness/lm_eval/api/model.py 72f5b387b64e9d9b |
unverified |
MIT (permissive) |
| Language Models are Few-Shot Learners |
28 May 2020 |
juletx/lm-evaluation-harness/lm_eval/api/model.py 3e9d38ecacd03fa7 |
unverified |
MIT (permissive) |
| arXiv:openreview_QrC8OgQyOI |
|
viiika/Prism/LLaDA/LLaDA_Prism/dllm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:2025.naacl-long.237 |
|
nbasyl/LLM-FP4/lm_eval/base.py ea06eaae4fc1eaf0 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| arXiv:2025.findings-acl.744 |
|
szu-tera/RankedVotingSC/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| arXiv:2024.findings-emnlp.86 |
|
FloatAI/TKEval/evaluation/lm-evaluation-harness/lm_eval/api/model.py 20a7cc804eb22661 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |