| LT2: Linear-Time Looped Transformers added by Syntology |
2026-05 (from id) |
chili-lab/LT2/lingua/args.py 15be2e56e073a228 |
ran
|
BSD-3-Clause (permissive) |
| Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization added by Syntology |
2026-05 (from id) |
Hyalinesky/DGAO/trl/core.py dfd0ecd9cb25af69 |
ran
|
Apache-2.0 (permissive) |
| UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing added by Syntology |
2026-04 (from id) |
Yunkaidang/UHR/with-SAM/longva/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm added by Syntology |
2026-02 (from id) |
zjr2000/SPES/spes/safetensors_util.py b593a4bbddcc66c3 |
unverified |
Apache-2.0 (permissive) |
| Altruism and Fair Objective in Mixed-Motive Markov games added by Syntology |
2026-02 (from id) |
AkuBrains/altruistic-fair-MARL/src/train_suite.py 925f85148db8b0fe |
unverified |
Apache-2.0 (permissive) |
| VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding added by Syntology |
2026-01 (from id) |
ziHoHe/VidLaDA/train/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Cross-Layer Injection for Deep Vision-Language Fusion added by Syntology |
2026-01 (from id) |
codefuse-ai/CLI/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Self-Supervised Dynamical System Representations for Physiological Time-Series added by Syntology |
2025-12 (from id) |
yuqinie98/PatchTST/PatchTST_self_supervised/src/utils.py 4ac8e022e2aa0b43 |
unverified |
Apache-2.0 (permissive) |
| Equivariance by Contrast: Identifiable Equivariant Embeddings from Unlabeled Finite Group Actions added by Syntology |
2025-10 (from id) |
dynamical-inference/ebc/groupcl/datajoint/sweep.py f30ac4807ed6c8c5 |
unverified |
no licence file found · pointer only |
| Iterative Prompt Refinement for Safer Text-to-Image Generation added by Syntology |
2025-09 (from id) |
ku-dmlab/IPR/trl/core.py dfd0ecd9cb25af69 |
ran
|
Apache-2.0 (permissive) |
| Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time added by Syntology |
2025-09 (from id) |
Yifan-Lan/Phi/trl/core.py a51ae45a1f6c0b5b |
unverified |
MIT (permissive) |
| arXiv:2507.11630 |
2025-07 (from id) |
AlignmentResearch/harmtune/harmtune/utils/utils.py 6784396112f94bd1 |
unverified |
no licence file found · pointer only |
| From Bytes to Ideas: Language Modeling with Autoregressive U-Nets |
17 Jun 2025 |
facebookresearch/lingua/lingua/args.py 15be2e56e073a228 |
ran
|
BSD-3-Clause (permissive) |
| Steering LLM Thinking with Budget Guidance |
16 Jun 2025 |
umass-embodied-agi/budgetguidance/3rdparty/trl/trl/core.py dfd0ecd9cb25af69 |
ran
|
MIT (permissive) |
| Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors |
30 May 2025 |
LaVi-Lab/Video-3D-LLM/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| LaViDa: A Large Diffusion Language Model for Multimodal Understanding |
22 May 2025 |
jacklishufan/lavida/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen! |
21 May 2025 |
thu-coai/backdoor-data-extraction/train/trl/core.py dfd0ecd9cb25af69 |
ran
|
MIT (permissive) |
| How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study |
21 May 2025 |
thu-coai/lrm-safety-study/trl/trl/core.py dfd0ecd9cb25af69 |
ran
|
MIT (permissive) |
| DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models |
14 May 2025 |
xiwenc1/dra-grpo/src/open_r1/trl/core.py dfd0ecd9cb25af69 |
ran
|
MIT (permissive) |
| Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models |
28 Apr 2025 |
minnesotanlp/mpo/trl/core.py dfd0ecd9cb25af69 |
ran
|
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Reasoning to Learn from Latent Thoughts |
24 Mar 2025 |
ryoungj/BoLT/lingua/args.py 15be2e56e073a228 |
ran
|
Apache-2.0 (permissive) |
| Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment |
16 Jan 2025 |
tatsu-lab/alpaca_farm/src/alpaca_farm/common.py f5ff77db7757bc92 |
ran
|
Apache-2.0 (permissive) |
| Risk-Averse Finetuning of Large Language Models |
12 Jan 2025 |
sapanachaudhary/ra-rlhf/trl/core.py 1782ce5551146615 |
unverified |
Apache-2.0 (permissive) |
| 2 OLMo 2 Furious |
31 Dec 2024 |
allenai/olmo/olmo/safetensors_util.py b593a4bbddcc66c3 |
unverified |
Apache-2.0 (permissive) |
| Memory Layers at Scale |
12 Dec 2024 |
facebookresearch/memory/lingua/args.py 15be2e56e073a228 |
ran
|
no licence file found · pointer only |
| DriveMM: All-in-One Large Multimodal Model for Autonomous Driving |
10 Dec 2024 |
zhijian11/DriveMM/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning |
4 Dec 2024 |
lavi-lab/aim/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models |
22 Nov 2024 |
kd-tao/dycoke/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization |
17 Nov 2024 |
APiaoG/SymDPO/trl/core.py a51ae45a1f6c0b5b |
unverified |
no licence file found · pointer only |
| CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset |
18 Nov 2024 |
meize0729/ccexpert/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Generating Highly Designable Proteins with Geometric Algebra Flow Matching |
7 Nov 2024 |
hits-mli/gafl/gafl/experiment_utils.py 262dce4f87fc7ad0 |
unverified |
licence not identified · pointer only |
| MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models |
23 Oct 2024 |
liuziyu77/mia-dpo/LLaVA-Hound-DPO/llava_hound_dpo/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/trl/core.py a51ae45a1f6c0b5b |
unverified |
no licence file found · pointer only |
| IntersectionZoo: Eco-driving for Benchmarking Multi-Agent Contextual Reinforcement Learning |
19 Oct 2024 |
mit-wu-lab/IntersectionZoo/code/policy_evaluation.py 15be2e56e073a228 |
ran
|
MIT (permissive) |
| Self-supervised contrastive learning performs non-linear system identification |
18 Oct 2024 |
dynamical-inference/dcl/dcl/datajoint/sweep.py f30ac4807ed6c8c5 |
unverified |
Apache-2.0 (permissive) |
| Do Unlearning Methods Remove Information from Language Model Weights? |
11 Oct 2024 |
aghyad-deeb/unlearning_evaluation/pipeline.py 0260364dcc364961 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning |
8 Oct 2024 |
Harry67Hu/CORY/trl/core.py a51ae45a1f6c0b5b |
unverified |
MIT (permissive) |
| Temperature Optimization for Bayesian Deep Learning |
8 Oct 2024 |
weiyaw/tempered-posteriors/visual_utils.py 91345e63f4999e0e |
unverified |
MIT (permissive) |
| Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition |
2 Oct 2024 |
kaistAI/Knowledge-Entropy/olmo/safetensors_util.py b593a4bbddcc66c3 |
unverified |
Apache-2.0 (permissive) |
| CableInspect-AD: An Expert-Annotated Anomaly Detection Dataset |
30 Sep 2024 |
mila-iqia/cableinspect-ad-code/src/enhanced-patchcore/post_processing/post_processing_unsupervised_results.py 94bd9041d52daa8d |
ran
|
Apache-2.0 (permissive) |
| Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling |
30 Aug 2024 |
jubayer-hamid/bid_lerobot/lerobot/common/datasets/utils.py e2adf71cdd509f3e |
ran
|
Apache-2.0 (permissive) |
| xLSTMTime : Long-term Time Series Forecasting With xLSTM |
14 Jul 2024 |
muslehal/xLSTMTime/src/utils.py 4ac8e022e2aa0b43 |
unverified |
MIT (permissive) |
| Suri: Multi-constraint Instruction Following for Long-form Text Generation |
27 Jun 2024 |
chtmp223/suri/ft/lib/trl_mod/core.py a51ae45a1f6c0b5b |
unverified |
no licence file found · pointer only |
| ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO |
17 Jun 2024 |
yonseivnl/vlm-rlaif/RLAIF/data_utils/common_utils.py f5ff77db7757bc92 |
ran
|
Apache-2.0 (permissive) |
| ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO |
17 Jun 2024 |
snumprlab/SRT/trl/core.py a51ae45a1f6c0b5b |
unverified |
no licence file found · pointer only |
| On Subjective Uncertainty Quantification and Calibration in Natural Language Generation |
7 Jun 2024 |
meta-inf/suq-nlg/generation/utils.py 6afd679238dc532d |
ran
|
MIT (permissive) |
| Efficient Adversarial Training in LLMs with Continuous Attacks |
24 May 2024 |
sophie-xhonneux/continuous-advtrain/src/database_handling.py 15be2e56e073a228 |
ran
|
MIT (permissive) |
| ALaRM: Align Language Models via Hierarchical Rewards Modeling |
11 Mar 2024 |
halfrot/ALaRM/trl/trl/core.py 1782ce5551146615 |
unverified |
Apache-2.0 (permissive) |
| Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models |
28 Feb 2024 |
mignonjia/ts_watermark/utils/submitit.py 55a42447218367a6 |
ran
|
no licence file found · pointer only |
| Model Editing by Standard Fine-Tuning |
16 Feb 2024 |
au-revoir/model-editing-ft/trl/core.py 1782ce5551146615 |
unverified |
no licence file found · pointer only |
| Personalized Language Modeling from Personalized Human Feedback |
6 Feb 2024 |
humainlab/personalized_rlhf/evaluate/alpaca_farm/common.py f5ff77db7757bc92 |
ran
|
MIT (permissive) |
| Neural General Circulation Models for Weather and Climate |
13 Nov 2023 |
neuralgcm/neuralgcm/neuralgcm/reference_code/train_utils.py 4e71b8bf4cde8432 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Diversify Question Generation with Retrieval-Augmented Style Transfer |
23 Oct 2023 |
gouqi666/RAST/rag/core.py ddf7639e1bf7edf8 |
unverified |
no licence file found · pointer only |
| Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging |
17 Oct 2023 |
joeljang/rlphf/gpt4_evaluate/alpaca_farm/common.py f5ff77db7757bc92 |
ran
|
no licence file found · pointer only |
| GRANDE: Gradient-Based Decision Tree Ensembles for Tabular Data |
29 Sep 2023 |
s-marton/gradtree/experiments_paper_gradtree/utilities/utilities_GDT.py 5cd213cdf1030eee |
ran
|
MIT (permissive) |
| Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF |
16 Sep 2023 |
simengsun/alpaca_farm_lora/alpaca_farm/src/alpaca_farm/common.py f5ff77db7757bc92 |
ran
|
Apache-2.0 (permissive) |
| In-Context Learning Learns Label Relationships but Is Not Conventional Learning |
23 Jul 2023 |
jlko/in_context_learning/llm/plot_utils.py 8e87d24bb6fec877 |
ran
|
no licence file found · pointer only |
| Locally Adaptive Federated Learning |
12 Jul 2023 |
IssamLaradji/sps/src/utils.py 0e62b13e0a76c17b |
ran
|
no licence file found · pointer only |
| Direct Preference Optimization: Your Language Model is Secretly a Reward Model |
29 May 2023 |
padlex/trl/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| Distance Weighted Supervised Learning for Offline Interaction Data |
26 Apr 2023 |
jhejna/dwsl/research/algs/dwsl.py 0954fb2e0809ee19 |
ran · our draft was wrong
|
MIT (permissive) |
| On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting |
1 Jun 2022 |
naver/gdc/cdpg/cdpg/core.py 8826b925c6d6132e |
unverified |
licence not identified · pointer only |
| Manifold Characteristics That Predict Downstream Task Performance |
16 May 2022 |
bytefuse/representation-manifold-quality-metric/src/utils.py ab7cb2f6ec7f6e67 |
unverified |
MIT (permissive) |
| ChaCha for Online AutoML |
9 Jun 2021 |
microsoft/FLAML/flaml/tune/scheduler/online_scheduler.py fbb7d311c959569d |
ran · our draft was wrong
|
MIT (permissive) |
| Causally motivated Shortcut Removal Using Auxiliary Labels |
13 May 2021 |
mymakar/causally_motivated_shortcut_removal/shared/train_utils.py 3e2e16445f6f92e6 |
unverified |
Apache-2.0 (permissive) |
| Deep Learning Hamiltonian Monte Carlo |
7 May 2021 |
saforem2/l2hmc-qcd/src/l2hmc/configs.py 3e2181383d727aba |
unverified |
Apache-2.0 (permissive) |
| How Sensitive are Meta-Learners to Dataset Imbalance? |
12 Apr 2021 |
mattochal/imbalanced_fsl_public/generator.py 0b2f8809055e38f5 |
ran · our draft was wrong
|
MIT (permissive) |
| Few-Shot Learning with Class Imbalance |
7 Jan 2021 |
identical code first harvested elsewhere 0b2f8809055e38f5 |
ran · our draft was wrong
|
licence of this copy not recorded |
| A Distributional Approach to Controlled Text Generation |
21 Dec 2020 |
naver/gdc/dpg/gdc/gdc/core.py 2f2b2a8fe27dcc79 |
unverified |
licence not identified · pointer only |
| A Distributional Approach to Controlled Text Generation |
21 Dec 2020 |
identical code first harvested elsewhere 8826b925c6d6132e |
unverified |
licence of this copy not recorded |
| Generalized Negative Correlation Learning for Deep Ensembling |
5 Nov 2020 |
sbuschjaeger/Pysembles/pysembles/Utils.py 9920339eff7f6691 |
ran · our draft was wrong
|
MIT (permissive) |
| Global optimization of Lipschitz functions |
7 Mar 2017 |
Sycor4x/lipo/benchmarks.py 9bbabe0c6fef21bb |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| arXiv:Zhong_AIM_Adaptive_Inference_of_Multi-Modal_LLMs_via_Token_Merging_and_ICCV_2025_paper |
|
LaVi-Lab/AIM/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.findings-acl.359 |
|
G-JWLee/TAMP/trl/core.py a51ae45a1f6c0b5b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.acl-long.732 |
|
wchrepo/mulfe/evaluation/utils.py fc1cb4f8e17c39aa |
unverified |
Apache-2.0 (permissive) |