| ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation added by Syntology |
2026-09 (from id) |
speridlabs/eneas/eneas/vendor/SeC/inference/modeling_phi3.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation added by Syntology |
2026-09 (from id) |
speridlabs/eneas/eneas/vendor/SeC/inference/modeling_internlm2.py 051ebcfa8673a02e |
unverified |
Apache-2.0 (permissive) |
| Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding added by Syntology |
2026-09 (from id) |
qluoluo/faster-flash-decoding/ffd_core/modeling/llama.py de11bf4368235366 |
unverified |
Apache-2.0 (permissive) |
| FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation added by Syntology |
2026-08 (from id) |
FireRedTeam/FireRedAudio/fireredaudio/flow/x_transformers.py 587e69a3ace67351 |
ran
|
Apache-2.0 (permissive) |
| Decoupled Vision-Language System for Multimodal Understanding and Generation added by Syntology |
2026-08 (from id) |
YifanXu74/Libra/libra/models/libra/modeling_libra.py 45b8a2c070efc3ed |
ran
|
Apache-2.0 (permissive) |
| RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation added by Syntology |
2026-08 (from id) |
RoyChao19477/RT-SEMamba/models/transformer_block.py 3a782e0711823a97 |
unverified |
licence not identified · pointer only |
| Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization added by Syntology |
2026-07 (from id) |
ShanechiLab/WiCAT/wicat/models/transformer.py 4fd6fbd2e5d5a63e |
ran
fingerprinted |
licence not identified · pointer only |
| Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS-ANS Dynamics added by Syntology |
2026-07 (from id) |
AutoBrain-sleep/OmniSleep/src/omnisleep/models.py 0f92704bc4a6e1f6 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based Recommenders added by Syntology |
2026-06 (from id) |
Tiny-Snow/GenRec/src/genrec/models/model_seqrec/sasrec_sprint.py b715a071c5e5950e |
ran · our draft was wrong
|
GPL-3.0 (copyleft) · pointer only |
| Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation added by Syntology |
2026-06 (from id) |
PunchlineAAAA/DICE/dream/modeling_dream.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment added by Syntology |
2026-06 (from id) |
facebookresearch/esm/esm/rotary_embedding.py 2a127d3ae372d7bd |
ran
fingerprinted |
MIT (permissive) |
| JETSPEC: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting added by Syntology |
2026-06 (from id) |
hao-ai-lab/JetSpec/jetspec/models/draft_head.py 517fb615df4dde0c |
unverified |
MIT (permissive) |
| REPAIR: Predictive Self-Supervised Representation Learning in Chess added by Syntology |
2026-06 (from id) |
Artificial-Chrisi/RePAIR/networks.py 64e4135bdebae85b |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models added by Syntology |
2026-06 (from id) |
llaraspata/HallucinationDetection/src/model/LlaMa.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings added by Syntology |
2026-06 (from id) |
CentreChen/EmbFilter/models/modeling_llama_lr.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings added by Syntology |
2026-06 (from id) |
CentreChen/EmbFilter/models/modeling_mistral_lr.py 583539efd6fd01fb |
ran · our draft was wrong
|
MIT (permissive) |
| IR3DE: A Linear Router for Large Language Models added by Syntology |
2026-06 (from id) |
gensyn-ai/IR3DE/models/modeling_llama.py 583539efd6fd01fb |
ran · our draft was wrong
|
MIT (permissive) |
| TreeFlash: Parallel AR-Approximation for Faster Speculative Decoding added by Syntology |
2026-06 (from id) |
ETH-DISCO/TreeFlash/tree_flash.py cce1f5188f2264e4 |
unverified |
no licence file found · pointer only |
| DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation added by Syntology |
2026-06 (from id) |
SAI-Lab-NYU/DREAM-S/dream_s/model/cnets.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
no licence file found · pointer only |
| DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation added by Syntology |
2026-06 (from id) |
SAI-Lab-NYU/DREAM-S/dream_s/model/modeling_llama_kv.py bad260f6d71db00c |
unverified |
no licence file found · pointer only |
| Eigenvectors of Experts are Training-free Non-collapsing Routers added by Syntology |
2026-05 (from id) |
giangdip2410/SSMoE/SSMoE/Language/SSMoE-Embedding/models/modeling_deepseek.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Eigenvectors of Experts are Training-free Non-collapsing Routers added by Syntology |
2026-05 (from id) |
giangdip2410/SSMoE/SSMoE/Language/SSMoE-Embedding/models/modeling_olmoe.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Periodic RoPE for Infinite Context LLMs added by Syntology |
2026-05 (from id) |
Cominder/miniwin/model/model_miniwin.py 41e3bef5aa57fbfa |
unverified |
Apache-2.0 (permissive) |
| One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs added by Syntology |
2026-05 (from id) |
hed-ucas/Layer-wise-Learning-Rate/galore_utils/modeling_llama.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning added by Syntology |
2026-05 (from id) |
misonsky/chunk/models/modeling_llama.py 0fe82a947dc39b42 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era added by Syntology |
2026-05 (from id) |
lluosi/RDB-CL/ETrain/Models/Qwen/modeling_qwen.py fbda921818ec02df |
unverified |
no licence file found · pointer only |
| One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer added by Syntology |
2026-05 (from id) |
juchengshen/air/models/air/air_1net_L2x_H2x_input_token_prepend.py 99138bc8ae291d5f |
unverified |
Apache-2.0 (permissive) |
| Mela: Test-Time Memory Consolidation based on Transformation Hypothesis added by Syntology |
2026-05 (from id) |
Musubi-ai/Mela/mela/model/commons.py 583539efd6fd01fb |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment added by Syntology |
2026-05 (from id) |
piooip/SynerMedGen/modeling/bagel/siglip_navit.py 4102c144078044c7 |
ran
|
Apache-2.0 (permissive) |
| TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis added by Syntology |
2026-04 (from id) |
xiaomi-research/tts-prism/models/mimo_audio_tokenizer/modeling_rope_utils.py 80d3a54eab8134eb |
unverified |
Apache-2.0 (permissive) |
| GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion added by Syntology |
2026-04 (from id) |
mikumifa/GS-Quant/codebook/rqvae.py b0bb9cd4b9fd6357 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| R 2 -dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction added by Syntology |
2026-04 (from id) |
GATECH-EIC/R2-dLLM/dream/model/modeling_dream.py ac6507e437413e71 |
ran
|
no licence file found · pointer only |
| DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference added by Syntology |
2026-04 (from id) |
X-Xia0828/DepCap/dream/model/modeling_dream.py ac6507e437413e71 |
ran
|
no licence file found · pointer only |
| Gumbel Distillation for Parallel Text Generation added by Syntology |
2026-03 (from id) |
hxixixh/gumbel-distill/models/dit_gumbel.py 2700f6838f384546 |
unverified |
no licence file found · pointer only |
| Learning Transferable Sensor Models via Language-Informed Pretraining added by Syntology |
2026-03 (from id) |
yuc0805/SLIP/modeling_slip.py 6ab251ef16ca3891 |
unverified |
no licence file found · pointer only |
| SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization SERQ: SALIENCY-AWARE LOW-RANK ERROR RECONSTRUCTION FOR LLM QUANTIZATION added by Syntology |
2026-03 (from id) |
acalabys/SERQ/modeling/modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization SERQ: SALIENCY-AWARE LOW-RANK ERROR RECONSTRUCTION FOR LLM QUANTIZATION added by Syntology |
2026-03 (from id) |
acalabys/SERQ/modeling/modeling_gemma3.py 583539efd6fd01fb |
ran · our draft was wrong
|
no licence file found · pointer only |
| SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training added by Syntology |
2026-03 (from id) |
swamynathanvp/Surprisal-Aware-Residual-Test-Time-Training/sr_ttt_fixed.py e96c58f6ba701d7c |
unverified |
MIT (permissive) |
| SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training added by Syntology |
2026-03 (from id) |
swamynathanvp/Surprisal-Aware-Residual-Test-Time-Training/sr_ttt_kaggle.py db4420e92850c6d0 |
unverified |
MIT (permissive) |
| SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training added by Syntology |
2026-03 (from id) |
swamynathanvp/Surprisal-Aware-Residual-Test-Time-Training/sr_ttt_v2.py 42256aaa3ae587e2 |
unverified |
MIT (permissive) |
| Stacked from One: Multi-Scale Self-Injection for Context Window Extension added by Syntology |
2026-03 (from id) |
Clement25/SharedLLM/modeling/modeling_llama_sharedllm_flash.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement added by Syntology |
26 Feb 2026 |
SHUCHENZHU/LITE/peft_pretraining/modeling_llama.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Large Causal Models for Temporal Causal Discovery added by Syntology |
2026-02 (from id) |
kougioulis/LCM/lcm_pytorch/models/misc/RotaryEmbedding.py 045feff71ed98d11 |
unverified |
Apache-2.0 (permissive) |
| MoE-Spec: Expert Budgeting for Efficient Speculative Decoding added by Syntology |
2026-02 (from id) |
SafeAILab/EAGLE/eagle/modeling_eagle.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Doc-to-LoRA: Learning to Instantly Internalize Contexts added by Syntology |
2026-02 (from id) |
SakanaAI/doc-to-lora/src/ctx_to_lora/modeling/text_to_lora.py 22f75dc2bea7b46e |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| DAWN: Dependency-Aware Fast Inference for Diffusion LLMs added by Syntology |
2026-02 (from id) |
lizhuo-luo/DAWN/dream/model/modeling_dream.py ac6507e437413e71 |
ran
|
MIT (permissive) |
| CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs added by Syntology |
2026-02 (from id) |
hrlics/CoPE/modeling_cope.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Audio ControlNet for Fine-Grained Audio Generation and Editing added by Syntology |
2026-02 (from id) |
haidog-yaqub/EzAudio/src/models/controlnet.py 206c90f5d5d477d1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| LLM4Fluid: Large Language Models as Generalizable Neural Solvers for Fluid Dynamics added by Syntology |
29 Jan 2026 |
BaratiLab/FactFormer/libs/factorization_module.py c3acddf93e491a49 |
ran · our draft was wrong
|
MIT (permissive) |
| APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention added by Syntology |
2026-01 (from id) |
thunlp/APB/experiments_apb/APB/modeling_llama_locret.py e12c6e1855e9fbe0 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference added by Syntology |
2026-01 (from id) |
fvliang/DART/dart/model/modeling_mixtral_kv.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference added by Syntology |
2026-01 (from id) |
fvliang/DART/dart/model/modeling_qwen3_kv.py 583539efd6fd01fb |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference added by Syntology |
2026-01 (from id) |
fvliang/DART/dart/model/llama3_dart.py a93e62ca7757414c |
unverified |
Apache-2.0 (permissive) |
| DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference added by Syntology |
2026-01 (from id) |
fvliang/DART/dart/model/modeling_llama_kv.py eee34413e2bccab6 |
unverified |
Apache-2.0 (permissive) |
| VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding added by Syntology |
2026-01 (from id) |
ziHoHe/VidLaDA/train/llada_v_prepare/files/modeling_llada.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers added by Syntology |
2026-01 (from id) |
LCM-Lab/Elastic-Attention/elasticattn/src/model_utils.py 583539efd6fd01fb |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification added by Syntology |
2026-01 (from id) |
5SSjw/MARS/mars/model/cnets1.py f725bc2d76076485 |
ran · fixture could not drive it
|
MIT (permissive) |
| MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification added by Syntology |
2026-01 (from id) |
5SSjw/MARS/mars/model/cnets.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
MIT (permissive) |
| MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification added by Syntology |
2026-01 (from id) |
5SSjw/MARS/mars/model/modeling_llama_kv.py eee34413e2bccab6 |
unverified |
MIT (permissive) |
| UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation added by Syntology |
2026-01 (from id) |
ZrH42/UniX/modeling/unix/siglip_navit.py 4102c144078044c7 |
ran
|
MIT (permissive) |
| STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning added by Syntology |
2026-01 (from id) |
LingFengGold/STReasoner/base_model/Config-Qwen2.5-14B-Instruct/modeling_qwen2.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| FOAM: Blocked State Folding for Memory-Efficient LLM Training added by Syntology |
2025-12 (from id) |
zqOuO/FOAM/peft_pretraining/modeling_llama.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Few-shot Protein Fitness Prediction via In-context Learning and Test-time Training added by Syntology |
2025-12 (from id) |
fteufel/PRIMO/primo/zero_shot_utils/modeling_progen.py 2f6618868885e7ca |
ran · fixture could not drive it
|
no licence file found · pointer only |
| LiveStar: Live Streaming Assistant for Real-World Online Video Understanding added by Syntology |
2025-11 (from id) |
yzy-bupt/LiveStar/inference/modeling_livestar_llm.py 051ebcfa8673a02e |
unverified |
no licence file found · pointer only |
| IG-Pruning: Input-Guided Block Pruning for Large Language Models added by Syntology |
2025-11 (from id) |
ictnlp/IG-Pruning/src/models/dsr_modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Information-Theoretic Discrete Diffusion added by Syntology |
2025-10 (from id) |
Dongjae0324/infodis/model/rotary.py b3c9c4aebf03fa97 |
unverified |
no licence file found · pointer only |
| Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs added by Syntology |
2025-10 (from id) |
LiuJinzhe-Keepgoing/NMKE/modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Sparser Block-Sparse Attention via Token Permutation added by Syntology |
2025-10 (from id) |
xinghaow99/pbs-attn/pbs_attn/patch/huggingface.py de11bf4368235366 |
unverified |
no licence file found · pointer only |
| MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization added by Syntology |
2025-10 (from id) |
pkumelon/MeCeFO/peft_pretraining/modeling_llama.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model added by Syntology |
2025-10 (from id) |
RuipingL/Situat3DChange/SCReasoner/model/transformers.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
CC-BY-4.0 · pointer only |
| Towards Generalizable PDE Dynamics Forecasting via Physics-Guided Invariant Learning added by Syntology |
2025-09 (from id) |
LSY-Cython/iMOOE/models/framework.py f1c3b8a106802591 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding added by Syntology |
2025-09 (from id) |
KangJialiang/ViSpec/vispec/model/modeling_mixtral_kv.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding added by Syntology |
2025-09 (from id) |
KangJialiang/ViSpec/vispec/model/cnets.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding added by Syntology |
2025-09 (from id) |
KangJialiang/ViSpec/vispec/model/modeling_llama_kv.py bad260f6d71db00c |
unverified |
Apache-2.0 (permissive) |
| Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision added by Syntology |
2025-08 (from id) |
Fr0zenCrane/UniCoT/modeling/qwen2/modeling_qwen2.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision added by Syntology |
2025-08 (from id) |
Fr0zenCrane/UniCoT/modeling/bagel/siglip_navit.py 4102c144078044c7 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2507.23284 |
2025-07 (from id) |
mlvlab/BLiM/videochat_flash/modeling_qwen2_flash.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:2507.22424 |
2025-07 (from id) |
PineTreeWss/SpecVLA/openvla/specdecoding/model/modeling_mixtral_kv.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:2507.22424 |
2025-07 (from id) |
PineTreeWss/SpecVLA/openvla/specdecoding/model/cnets.py f725bc2d76076485 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:2507.22424 |
2025-07 (from id) |
PineTreeWss/SpecVLA/openvla/specdecoding/model/modeling_llama_kv.py eee34413e2bccab6 |
unverified |
MIT (permissive) |
| arXiv:2507.17539 |
2025-07 (from id) |
MeteorElf/FundusExpert/src/internvl_chat/internvl/patch/llama2_flash_attn_monkey_patch.py 3408dd47e939d760 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.10419 |
2025-07 (from id) |
Victorletzelter/LoRA-MCL/ALMA/modeling_xalma.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2507.01299 |
2025-07 (from id) |
alibaba/EfficientAI/larosa/inference/modeling_llama_larosa.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models |
25 Jun 2025 |
thecommonirin/qresafe/quant-with-ft/models/modeling_llama_quant.py f725bc2d76076485 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning |
23 Jun 2025 |
thudm/longwriter/train/patch/modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning |
23 Jun 2025 |
thudm/longwriter/train/patch/modeling_chatglm.py 5cfeaee89a75c5ff |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Show-o2: Improved Native Unified Multimodal Models |
18 Jun 2025 |
showlab/Show-o/show-o2/models/modules.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking |
11 Jun 2025 |
princeton-pli/QRHead/src/qrretriever/custom_modeling_qwen2.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking |
11 Jun 2025 |
princeton-pli/QRHead/src/qrretriever/custom_modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/models/llama/modeling_llama.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching |
17 May 2025 |
maomaocun/dLLM-cache/dllm_cache/hooks/cache_hook_LLaDA_V.py ebbbedae1f60f76d |
unverified |
Apache-2.0 (permissive) |
| Pretraining Language Models to Ponder in Continuous Space |
27 May 2025 |
identical code first harvested elsewhere bac65c3dafaec040 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation |
24 May 2025 |
ashaba1in/smoothie/model/score_estimator.py a4f77e43ad34dc66 |
unverified |
no licence file found · pointer only |
| Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation |
22 May 2025 |
ruizheliUOA/ARC_JSD/llm_models/modeling_gemma2.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning |
22 May 2025 |
jiaruzouu/transformercopilot/src/Decoder_Only/copilot/modeling_flax_llama.py a1b46b3f84e76f6a |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Emerging Properties in Unified Multimodal Pretraining |
20 May 2025 |
neverbiasu/ComfyUI-BAGEL/modeling/bagel/siglip_navit.py 4102c144078044c7 |
ran
|
Apache-2.0 (permissive) |
| UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens |
20 May 2025 |
arctanxarc/unictokens/models/phi.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation |
20 May 2025 |
opengvlab/videochat-flash/llava-train_videochat/llava/model/language_model/modeling_qwen2_flash.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| Multi-Token Prediction Needs Registers |
15 May 2025 |
nasosger/mutor/language_modeling/src/models/llama_mutor.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| Multi-Token Prediction Needs Registers |
15 May 2025 |
nasosger/mutor/language_modeling/src/models/gemma_mutor.py 583539efd6fd01fb |
ran · our draft was wrong
|
MIT (permissive) |
| Parallel Scaling Law for Language Models |
15 May 2025 |
identical code first harvested elsewhere 583539efd6fd01fb |
ran · our draft was wrong
|
licence of this copy not recorded |
| Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free |
10 May 2025 |
qiuzh20/gated_attention/modeling_qwen3.py bac65c3dafaec040 |
ran · our draft was wrong
|
MIT (permissive) |
| RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale |
5 May 2025 |
recursal/radlads-paper/rwkv6qwen2/modeling_rwkv6qwen2.py bac65c3dafaec040 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale |
5 May 2025 |
recursal/radlads-paper/models/qwen2.py 67e215883a9a54b6 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2504.21403 |
2025-04 (from id) |
ANDgate99/Explore-Then-Select/model/modeling_qwen2.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| ReasonIR: Training Retrievers for Reasoning Tasks |
29 Apr 2025 |
identical code first harvested elsewhere bac65c3dafaec040 |
ran · our draft was wrong
|
licence of this copy not recorded |
| TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos |
24 Apr 2025 |
renshuhuai-andy/timechat/timechat/models/modeling_llama.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
BSD-3-Clause (permissive) |
| Why do LLMs attend to the first token? |
3 Apr 2025 |
zzmtsvv/ad-gta/src/nn/memeff_rope_fn.py e8f6d785640a6346 |
unverified |
Apache-2.0 (permissive) |
| Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequence |
26 Mar 2025 |
identical code first harvested elsewhere 0fe82a947dc39b42 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| Scaling Vision Pre-Training to 4K Resolution |
25 Mar 2025 |
efficient-large-model/vila/llava/eval/vision_niah_vila/zigzag_ring_attn/modeling_qwen2.py 5d81bffdb6022427 |
ran
|
Apache-2.0 (permissive) |
| Mixture of Lookup Experts |
20 Mar 2025 |
jieshibo/mole/modeling_mole.py fc2a2c4d5dc6bf94 |
unverified |
MIT (permissive) |
| MP-GUI: Modality Perception with MLLMs for GUI Understanding |
18 Mar 2025 |
BigTaige/MP-GUI/model/internvl/patch/llama2_flash_attn_monkey_patch.py 3408dd47e939d760 |
unverified |
MIT (permissive) |
| SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression |
16 Mar 2025 |
aiot-mlsys-lab/svd-llm/SVDLLM.py af9e30f5a1fe133e |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| FastVID: Dynamic Density Pruning for Fast Video Large Language Models |
14 Mar 2025 |
identical code first harvested elsewhere bac65c3dafaec040 |
ran · our draft was wrong
|
licence of this copy not recorded |
| TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models |
13 Mar 2025 |
shawntan86/tokencarve/TokenCarve/modeling_llama.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding |
13 Mar 2025 |
AMD-AIG-AIMA/Gumiho/gumiho/model/cnets.py 9a65a30d006fc96e |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding |
13 Mar 2025 |
AMD-AIG-AIMA/Gumiho/gumiho/model/modeling_llama_kv.py bad260f6d71db00c |
unverified |
Apache-2.0 (permissive) |
| Language-Enhanced Representation Learning for Single-Cell Transcriptomics |
12 Mar 2025 |
identical code first harvested elsewhere 9b4dff79d5e6102c |
ran · fixture could not drive it
|
licence of this copy not recorded |
| Implicit Reasoning in Transformers is Reasoning through Shortcuts |
10 Mar 2025 |
TianheL/LM-Implicit-Reasoning/src/model/modeling_gpt2_rope.py be41b9c696d434d2 |
ran
fingerprinted |
MIT (permissive) |
| ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration |
10 Mar 2025 |
idea-isail-lab-uiuc/resmoe/mixtral/resmoe_mixtral/modeling_mixtral.py d61c483a3c2b3156 |
ran · fixture could not drive it
|
MIT (permissive) |
| Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas |
3 Mar 2025 |
shiqichen17/AdaptVis/model_zoo/llama/modeling_llama.py 373a7df152e5a1a5 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference |
28 Feb 2025 |
bytedance/FlexPrefill/flex_prefill/modules/glm/glm_self_attention_foward.py 059a805ecbeca6af |
ran
|
Apache-2.0 (permissive) |