| Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture added by Syntology |
2026-08 (from id) |
fmarrabal/miuracognitive/model/transformer.py d388fa934b5e3f96 |
ran
|
licence not identified · pointer only |
| Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall A Pre-Registered Negative Result, with a Systems Measurement Study added by Syntology |
2026-08 (from id) |
Shriniwas410/cacheable-by-design/model.py 07941306772c339b |
unverified |
MIT (permissive) |
| Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion added by Syntology |
2026-08 (from id) |
ericu9500/stoicheia/hf_release/modeling_char_bert_joint.py 7b0a953e8c8431de |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models added by Syntology |
2026-07 (from id) |
SeunggeunKimkr/genuine-any-order/LatentMDM/model/latent_mdm.py bc605409c93e760b |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| FastMix: Fast Data Mixture Optimization via Gradient Descent added by Syntology |
2026-06 (from id) |
hrtan/fastmix/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment added by Syntology |
2026-05 (from id) |
black-forest-labs/flux/src/flux/math.py 3444bbcc9065f707 |
unverified |
Apache-2.0 (permissive) |
| CHESS-WORLD-MODEL: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences added by Syntology |
2026-05 (from id) |
Benjamin-Walker/Chess-World-Model/models/transformer.py 289ce4c835947e59 |
ran
fingerprinted |
MIT (permissive) |
| PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design added by Syntology |
2026-05 (from id) |
wang-henry4/prism/prism/model/prefix_material_thk_model.py 9c6b2f793d0099d9 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| No Free Swap: Protocol-Dependent Layer Redundancy in Transformers added by Syntology |
2026-05 (from id) |
Gpgabriel25/ProtocolGapDiagnostic/clean_oracle_bootstrap_ci.py 1d1f71ed01bdb63e |
unverified |
licence not identified · pointer only |
| parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning added by Syntology |
2026-05 (from id) |
xiaoyang-123-cell/ParallelCBF/parallelcbf/algorithms/causal_transformer.py 4fdbe3b55fbdf4b3 |
ran
fingerprinted |
licence not identified · pointer only |
| How to Scale Mixture-of-Experts: From µP to the Maximally Scale-Stable Parameterization added by Syntology |
2026-05 (from id) |
vankadara-lab/mssp-moe/transformer-moe-experiments/model.py 5d32ec6879733f85 |
ran · our draft was wrong
|
MIT (permissive) |
| VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use added by Syntology |
2026-05 (from id) |
vectrayx/vectrayx-nano-paper/training/transformer.py e4539c0c100abce2 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization added by Syntology |
2026-05 (from id) |
YupengSu/MuonQ/src/models.py 25e9d5a6a1f48f0f |
ran
|
Apache-2.0 (permissive) |
| ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation added by Syntology |
2026-01 (from id) |
Kwai-Kolors/ResTok/modeling/restok.py 3444bbcc9065f707 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.07101 |
2025-07 (from id) |
martin-marek/batch-size/finetuning/rope.py ae9c4ad2e718b1ea |
unverified |
MIT (permissive) |
| DataDecide: How to Predict Best Pretraining Data with Small Experiments |
15 Apr 2025 |
gair-nlp/prox/train/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| SkyLadder: Better and Faster Pretraining via Context Window Scheduling |
19 Mar 2025 |
sail-sg/skyladder/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Strategy Coopetition Explains the Emergence and Transience of In-Context Learning |
7 Mar 2025 |
identical code first harvested elsewhere 369d990a59246f7f |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| KV-Edit: Training-Free Image Editing for Precise Background Preservation |
24 Feb 2025 |
Xilluill/KV-Edit/models/kv_edit.py 7b46f68f6e5e0e3d |
unverified |
Apache-2.0 (permissive) |
| ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features |
6 Feb 2025 |
helblazer811/ConceptAttention/concept_attention/flux/dit_block.py 7f1b2506d32e33e3 |
unverified |
no licence file found · pointer only |
| LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer |
3 Feb 2025 |
showlab/LayerTracer/library/flux_models.py ad4327d0fd3beb3a |
unverified |
MIT (permissive) |
| FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing |
10 Dec 2024 |
holmesshuan/fireflow/src/flux/math.py 3444bbcc9065f707 |
unverified |
Apache-2.0 (permissive) |
| Selective Attention: Enhancing Transformer through Principled Context Control |
19 Nov 2024 |
umich-sota/selective_attention/lit_gpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| Taming Rectified Flow for Inversion and Editing |
7 Nov 2024 |
wangjiangshan0725/rf-solver-edit/FLUX_Image_Edit/src/flux/math.py 3444bbcc9065f707 |
unverified |
no licence file found · pointer only |
| Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities |
15 Oct 2024 |
gpt-omni/mini-omni2/litgpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| When Attention Sink Emerges in Language Models: An Empirical View |
14 Oct 2024 |
sail-sg/attention-sink/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining |
10 Oct 2024 |
beccabai/multi-agent-data-selection/src/lit_gpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
no licence file found · pointer only |
| Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale |
25 Sep 2024 |
identical code first harvested elsewhere 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| FLUX that Plays Music |
1 Sep 2024 |
feizc/fluxmusic/modules/layers.py 3444bbcc9065f707 |
unverified |
no licence file found · pointer only |
| Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming |
29 Aug 2024 |
gpt-omni/mini-omni/litgpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| 1.5-Pints Technical Report: Pretraining in Days, Not Months -- Your Language Model Thrives on Quality Data |
7 Aug 2024 |
Pints-AI/1.5-Pints/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs |
14 Jun 2024 |
ahans30/goldfish-loss/lit_gpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
Apache-2.0 (permissive) |
| MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models |
10 Jun 2024 |
cxcscmu/MATES/src/lit_gpt/model.py aa32c6faaa562b69 |
ran · violated contract
fingerprinted |
MIT (permissive) |
| MotionLLM: Understanding Human Behaviors from Human Motions and Videos |
30 May 2024 |
IDEA-Research/MotionLLM/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MotionLLM: Understanding Human Behaviors from Human Motions and Videos |
30 May 2024 |
IDEA-Research/MotionLLM/lit_llama/model.py ebb19c1f412b0e2a |
unverified |
licence not identified · pointer only |
| Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads |
30 May 2024 |
Avelina9X/direct-preference-heads/src/model/layers.py f7c918ee79c006bf |
unverified |
BSD-3-Clause (permissive) |
| What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation |
10 Apr 2024 |
identical code first harvested elsewhere 369d990a59246f7f |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition |
8 Feb 2024 |
Hypotheses-Paradise/UADF/lit_llama/model.py 7137105772c52074 |
unverified |
Apache-2.0 (permissive) |
| Distinguishing the Knowable from the Unknowable with Language Models |
5 Feb 2024 |
gahdritz/llm_uncertainty/lit-llama/model.py 16b3ba074f0d8074 |
unverified |
Apache-2.0 (permissive) |
| TinyLlama: An Open-Source Small Language Model |
4 Jan 2024 |
jzhang38/tinyllama/lit_gpt/model.py 5728c74084e12d35 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Universal Visual Decomposer: Long-Horizon Manipulation Made Easy |
12 Oct 2023 |
zcczhang/uvd/uvd/models/nn/transformer.py 7137105772c52074 |
unverified |
MIT (permissive) |
| Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition |
10 Oct 2023 |
srijith-rkr/whispering-llama/lit_llama/model.py 7137105772c52074 |
unverified |
MIT (permissive) |
| LLaMA: Open and Efficient Foundation Language Models |
27 Feb 2023 |
Lightning-AI/lit-llama/lit_llama/model.py 7137105772c52074 |
unverified |
Apache-2.0 (permissive) |
| Data Distributional Properties Drive Emergent In-Context Learning in Transformers |
22 Apr 2022 |
aadityasingh/icl-dynamics/models.py 369d990a59246f7f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |