| ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation added by Syntology |
2026-09 (from id) |
speridlabs/eneas/eneas/vendor/SeC/inference/modeling_internlm2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding added by Syntology |
2026-09 (from id) |
qluoluo/faster-flash-decoding/ffd_core/modeling/llama.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation added by Syntology |
2026-06 (from id) |
PunchlineAAAA/DICE/dream/modeling_dream.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding added by Syntology |
2026-06 (from id) |
TencentBAC/LazyMCoT/Qwen2.5/modeling_qwen2_5_vl_re_infer.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models added by Syntology |
2026-06 (from id) |
KQL11/Deep-VRM/Models/DeepVRM/modeling_qwen2_5_vl.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models added by Syntology |
2026-06 (from id) |
llaraspata/HallucinationDetection/src/model/LlaMa.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings added by Syntology |
2026-06 (from id) |
CentreChen/EmbFilter/models/modeling_llama_lr.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings added by Syntology |
2026-06 (from id) |
CentreChen/EmbFilter/models/modeling_mistral_lr.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| IR3DE: A Linear Router for Large Language Models added by Syntology |
2026-06 (from id) |
gensyn-ai/IR3DE/models/modeling_llama.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation added by Syntology |
2026-06 (from id) |
SAI-Lab-NYU/DREAM-S/dream_s/model/cnets.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation added by Syntology |
2026-06 (from id) |
SAI-Lab-NYU/DREAM-S/dream_s/model/modeling_llama_kv.py f6e0c0fc3868b787 |
ran
fingerprinted |
no licence file found · pointer only |
| Eigenvectors of Experts are Training-free Non-collapsing Routers added by Syntology |
2026-05 (from id) |
giangdip2410/SSMoE/SSMoE/Language/SSMoE-Embedding/models/modeling_deepseek.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Periodic RoPE for Infinite Context LLMs added by Syntology |
2026-05 (from id) |
Cominder/miniwin/model/model_miniwin.py 07d6ef500d143328 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning added by Syntology |
2026-05 (from id) |
misonsky/chunk/models/modeling_llama.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention added by Syntology |
2026-05 (from id) |
fasa-org/dash-attention/dash_attn/prefill/stage2_block_selection.py 4e762613ee907cdb |
ran · fixture could not drive it
fingerprinted |
BSD-3-Clause (permissive) |
| MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization added by Syntology |
2026-05 (from id) |
YupengSu/MuonQ/src/models.py 7140f5aa8cd118f3 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Mela: Test-Time Memory Consolidation based on Transformation Hypothesis added by Syntology |
2026-05 (from id) |
Musubi-ai/Mela/mela/model/modeling_mela.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention added by Syntology |
2026-05 (from id) |
flas-ai/FLAS/src/flas/model.py 0f07b44f9c18969e |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| R 2 -dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction added by Syntology |
2026-04 (from id) |
GATECH-EIC/R2-dLLM/dream/model/modeling_dream.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference added by Syntology |
2026-04 (from id) |
X-Xia0828/DepCap/dream/model/modeling_dream.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference added by Syntology |
2026-04 (from id) |
qqtang-code/FluxAttention/fluxattn/training/eval/modeling_flash_llama.py 56fbb035393380b6 |
unverified |
MIT (permissive) |
| SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization SERQ: SALIENCY-AWARE LOW-RANK ERROR RECONSTRUCTION FOR LLM QUANTIZATION added by Syntology |
2026-03 (from id) |
acalabys/SERQ/modeling/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization SERQ: SALIENCY-AWARE LOW-RANK ERROR RECONSTRUCTION FOR LLM QUANTIZATION added by Syntology |
2026-03 (from id) |
acalabys/SERQ/modeling/modeling_gemma3.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents added by Syntology |
2026-03 (from id) |
94wen94/ST-Lite/eval/attention_helpers.py 404c0af567476d50 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| MoE-Spec: Expert Budgeting for Efficient Speculative Decoding added by Syntology |
2026-02 (from id) |
SafeAILab/EAGLE/eagle/modeling_eagle.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| MoE-Spec: Expert Budgeting for Efficient Speculative Decoding added by Syntology |
2026-02 (from id) |
SafeAILab/EAGLE/eagle/model/cnets.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs added by Syntology |
2026-02 (from id) |
hrlics/CoPE/modeling_cope.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs added by Syntology |
2026-02 (from id) |
hrlics/CoPE/train/training/modeling_flash_llama_cope.py 69874376c94f0785 |
unverified |
no licence file found · pointer only |
| DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference added by Syntology |
2026-01 (from id) |
fvliang/DART/dart/model/llama3_dart.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding added by Syntology |
2026-01 (from id) |
ziHoHe/VidLaDA/train/llada_v_prepare/files/modeling_llada.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers added by Syntology |
2026-01 (from id) |
LCM-Lab/Elastic-Attention/elasticattn/src/model_utils.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification added by Syntology |
2026-01 (from id) |
5SSjw/MARS/mars/model/cnets1.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification added by Syntology |
2026-01 (from id) |
5SSjw/MARS/mars/model/cnets.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning added by Syntology |
2026-01 (from id) |
LingFengGold/STReasoner/base_model/Config-Qwen2.5-14B-Instruct/modeling_qwen2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| LiveStar: Live Streaming Assistant for Real-World Online Video Understanding added by Syntology |
2025-11 (from id) |
yzy-bupt/LiveStar/inference/modeling_livestar_llm.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| IG-Pruning: Input-Guided Block Pruning for Large Language Models added by Syntology |
2025-11 (from id) |
ictnlp/IG-Pruning/src/models/dsr_modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| IG-Pruning: Input-Guided Block Pruning for Large Language Models added by Syntology |
2025-11 (from id) |
ictnlp/IG-Pruning/src/models/dsr_modeling_qwen3.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs added by Syntology |
2025-10 (from id) |
LiuJinzhe-Keepgoing/NMKE/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Sparser Block-Sparse Attention via Token Permutation added by Syntology |
2025-10 (from id) |
xinghaow99/pbs-attn/pbs_attn/patch/huggingface.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Latent Speech-Text Transformer added by Syntology |
2025-10 (from id) |
facebookresearch/lst/lst/base_transformer.py 705c70c44ddc8080 |
unverified |
licence not identified · pointer only |
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding added by Syntology |
2025-09 (from id) |
KangJialiang/ViSpec/vispec/model/cnets.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding added by Syntology |
2025-09 (from id) |
KangJialiang/ViSpec/vispec/model/modeling_llama_kv.py f6e0c0fc3868b787 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision added by Syntology |
2025-08 (from id) |
Fr0zenCrane/UniCoT/modeling/qwen2/modeling_qwen2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:2507.23284 |
2025-07 (from id) |
mlvlab/BLiM/videochat_flash/modeling_qwen2_flash.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| arXiv:2507.22424 |
2025-07 (from id) |
PineTreeWss/SpecVLA/openvla/specdecoding/model/cnets.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding |
15 Jul 2025 |
ShiLuohe/KV-Latent/modelcodes/modeling_llama3.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| arXiv:2507.10419 |
2025-07 (from id) |
Victorletzelter/LoRA-MCL/ALMA/modeling_xalma.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning |
23 Jun 2025 |
thudm/longwriter/train/patch/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models |
18 Jun 2025 |
bytedance/video-salmonn-2/video_SALMONN2_pro/qwenvl/model/modeling_qwen3_vl.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking |
11 Jun 2025 |
princeton-pli/QRHead/src/qrretriever/custom_modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/models/llama/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Draft-based Approximate Inference for LLMs |
10 Jun 2025 |
identical code first harvested elsewhere 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs |
4 Jun 2025 |
the-scale-lab/AsymKV/asymkv/method/asymkv.py 56a1aae391d1d899 |
unverified |
Apache-2.0 (permissive) |
| Mamba Knockout for Unraveling Factual Information Flow |
30 May 2025 |
nirendy/mamba-knockout/src/experiments/knockout/llama/sdpa_attention.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction |
29 May 2025 |
identical code first harvested elsewhere 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation |
28 May 2025 |
arthursbianchessi/bam/models/bam.py 6d2a08dcf3466514 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Pretraining Language Models to Ponder in Continuous Space |
27 May 2025 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation |
22 May 2025 |
ruizheliUOA/ARC_JSD/llm_models/modeling_gemma2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens |
20 May 2025 |
arctanxarc/unictokens/models/phi.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation |
20 May 2025 |
opengvlab/videochat-flash/llava-train_videochat/llava/model/language_model/modeling_qwen2_flash.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning |
20 May 2025 |
jiwonsong-dev/reasoningpathcompression/rpc/rpc_utils.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Multi-Token Prediction Needs Registers |
15 May 2025 |
nasosger/mutor/language_modeling/src/models/llama_mutor.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Multi-Token Prediction Needs Registers |
15 May 2025 |
nasosger/mutor/language_modeling/src/models/gemma_mutor.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Parallel Scaling Law for Language Models |
15 May 2025 |
identical code first harvested elsewhere 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free |
10 May 2025 |
qiuzh20/gated_attention/modeling_qwen3.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale |
5 May 2025 |
recursal/radlads-paper/rwkv6attn.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:2504.21403 |
2025-04 (from id) |
ANDgate99/Explore-Then-Select/model/modeling_qwen2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| ReasonIR: Training Retrievers for Reasoning Tasks |
29 Apr 2025 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Efficient Reinforcement Finetuning via Adaptive Curriculum Learning |
7 Apr 2025 |
uscnlp-lime/verl/verl/models/transformers/monkey_patch.py 00f0f9848c8cd16c |
unverified |
Apache-2.0 (permissive) |
| Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequence |
26 Mar 2025 |
identical code first harvested elsewhere 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Scaling Vision Pre-Training to 4K Resolution |
25 Mar 2025 |
efficient-large-model/vila/llava/eval/vision_niah_vila/zigzag_ring_attn/modeling_qwen2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Mixture of Lookup Experts |
20 Mar 2025 |
jieshibo/mole/modeling_mole.py ab40781baaa35b67 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling |
17 Mar 2025 |
hustvl/MaTVLM/mamba2/hybrid_mamba_layer.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding |
16 Mar 2025 |
sczwangxiao/video-flexreduc/retake/longvideo_cache.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression |
16 Mar 2025 |
aiot-mlsys-lab/svd-llm/SVDLLM.py 25ef0712acc702e7 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| FastVID: Dynamic Density Pruning for Fast Video Large Language Models |
14 Mar 2025 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models |
13 Mar 2025 |
shawntan86/tokencarve/TokenCarve/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding |
13 Mar 2025 |
AMD-AIG-AIMA/Gumiho/gumiho/model/cnets.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding |
13 Mar 2025 |
AMD-AIG-AIMA/Gumiho/gumiho/model/modeling_llama_kv.py f6e0c0fc3868b787 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More |
17 Feb 2025 |
zichenwen1/dart/Qwen2_5-VL/Qwen2_5VL_DART/modeling_qwen2_5_vl_self.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Implicit Language Models are RNNs: Balancing Parallelization and Expressivity |
10 Feb 2025 |
microsoft/implicit_languagemodels/implicit_llm/modules/attention.py 541ba8da0f7b6512 |
unverified |
MIT (permissive) |
| LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models |
10 Feb 2025 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| VideoRoPE: What Makes for Good Video Rotary Position Embedding? |
7 Feb 2025 |
wiselnn570/videorope/videorope_plus/v_ruler/easy_context/modeling_qwen2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| M+: Extending MemoryLLM with Scalable Long-Term Memory |
1 Feb 2025 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting |
23 Jan 2025 |
brotherhappy/ostquant/models/modeling_llama.py ab40781baaa35b67 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing |
23 Jan 2025 |
mbzuai-oryx/geopixel/model/IXC/modeling_internlm2.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping |
11 Jan 2025 |
mayank31398/ladder-residual-inference/hf_modeling_utils/modeling_llama_ladder.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
BSD-3-Clause (permissive) |
| FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models |
30 Dec 2024 |
thu-nics/framefusion/framefusion/models/internvl/modeling_internlm2.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding |
29 Dec 2024 |
sczwangxiao/video-retake/retake/longvideo_cache.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding |
24 Dec 2024 |
cognitiveaisystems/3dgraphllm/models/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Experience of Training a 1.7B-Parameter LLaMa Model From Scratch |
17 Dec 2024 |
identical code first harvested elsewhere 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Byte Latent Transformer: Patches Scale Better Than Tokens |
13 Dec 2024 |
facebookresearch/blt/bytelatent/base_transformer.py 705c70c44ddc8080 |
unverified |
no licence file found · pointer only |
| Bayesian Optimization of Antibodies Informed by a Generative Model of Evolving Sequences |
10 Dec 2024 |
alannawzadamin/clonebo/clonebo/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Cross-Self KV Cache Pruning for Efficient Vision-Language Inference |
5 Dec 2024 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences |
2 Dec 2024 |
Hoyyyaard/LSceneLLM/models/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Scaling Particle Collision Data Analysis |
28 Nov 2024 |
supersymmetry-technologies/bbt-neutron/model/modeling_binbbt.py ab40781baaa35b67 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models |
7 Nov 2024 |
allenzren/open-pi-zero/src/model/utils.py 4e762613ee907cdb |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units |
4 Nov 2024 |
bkhmsi/llm-localization/models/modeling_gemma.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning |
25 Oct 2024 |
fyyfu/headkv/headkv/snapkv_utils.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing |
24 Oct 2024 |
yangyifei729/kvsharer/internlm2_real_share/modeling_internlm2_kvsharer.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Value Residual Learning For Alleviating Attention Concentration In Transformers |
23 Oct 2024 |
Zcchill/Value-Residual-Learning/src/modeling/modeling_llama_NeuTRENO_lambda04.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Scaling Diffusion Language Models via Adaptation from Autoregressive Models |
23 Oct 2024 |
HKUNLP/DiffuLLaMA/DiffuLLaMA-training/model_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction |
22 Oct 2024 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| MagicPIG: LSH Sampling for Efficient LLM Generation |
21 Oct 2024 |
infini-ai-lab/magicpig/models/utils.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Provable Benefits of Complex Parameterizations for Structured State Space Models |
17 Oct 2024 |
alxndrTL/mamba.py/mambapy/jamba.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| On the Role of Attention Heads in Large Language Model Safety |
17 Oct 2024 |
ydyjya/SafetyHeadAttribution/lib/utils/custommodel.py 3c76e52815c5401d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors |
16 Oct 2024 |
weixuan-wang123/SADI/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models |
15 Oct 2024 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| MoH: Multi-Head Attention as Mixture-of-Head Attention |
15 Oct 2024 |
identical code first harvested elsewhere 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Subspace Optimization for Large Language Models with Convergence Guarantees |
15 Oct 2024 |
pkumelon/Golore/zo-bench/modeling_llama.py 30d7eec482ebf6b1 |
ran · fixture could not drive it
fingerprinted |
GPL-3.0 (copyleft) · pointer only |