| IDEEA: training-free Input-Dependent stEEring via Activation cluster matching added by Syntology |
2026-09 (from id) |
DSL-Lab/IDEEA/common/modeling.py e3f80c2ee531022c |
unverified |
no licence file found · pointer only |
| StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions added by Syntology |
2026-09 (from id) |
Cha0Ga0/SWAPSTATE/src/hs_swap/models.py b88b122f84c5cb80 |
unverified |
no licence file found · pointer only |
| Latent Fact-Checking: Detecting Misinformation through Activation Engineering added by Syntology |
2026-08 (from id) |
Malta-Lab/LaFaCt/methods/base.py 6c2fa8a910a6e5fe |
unverified |
no licence file found · pointer only |
| When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO added by Syntology |
2026-08 (from id) |
CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO/grpo_trainer_utils.py 9ac32f5a2f8c738c |
unverified |
MIT (permissive) |
| Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making? added by Syntology |
2026-07 (from id) |
costa-nus/llm-emotion-decision/utils/hf.py 9bcca6a41327a4e1 |
unverified |
no licence file found · pointer only |
| Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression added by Syntology |
2026-06 (from id) |
Miraclecsc/Discrete-Latent-Reasoning/deepseek_codebook/visualize_latent_ids.py 8e067d12bb78f5b8 |
unverified |
no licence file found · pointer only |
| SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks added by Syntology |
2026-06 (from id) |
youai058/SlotGCG/baselines/model_utils.py 419da37a5e7e4f3a |
unverified |
no licence file found · pointer only |
| A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits added by Syntology |
2026-05 (from id) |
Yyuzrah/AMTA4R/MTA4R/src/eval_flexible.py 865ff8f8286bfd35 |
unverified |
no licence file found · pointer only |
| SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs added by Syntology |
2026-05 (from id) |
tally0818/SAGE/src/train/common.py f62f00e1e077e5a1 |
unverified |
no licence file found · pointer only |
| Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR added by Syntology |
2026-05 (from id) |
tally0818/NudgeRL/src/train/train_nudgerl.py 4238d2cfcafc21b0 |
ran
|
no licence file found · pointer only |
| Multi-Rollout On-Policy Distillation via Peer Successes and Failures added by Syntology |
2026-05 (from id) |
viviable/mopd_code/analysis/score_teacher_contexts.py 1a52c3aa3027b3fe |
unverified |
Apache-2.0 (permissive) |
| LLM-Agnostic Semantic Representation Attack added by Syntology |
2026-05 (from id) |
JiaweiLian/SRA/baselines/model_utils.py 090eed81a24d4b32 |
unverified |
MIT (permissive) |
| Latent Planning Emerges with Scale added by Syntology |
2026-04 (from id) |
hannamw/model-planning-public/animal_stories/probing.py fcdffa1e12fc52ec |
unverified |
MIT (permissive) |
| Same Geometry, Opposite Noise: Transformer Magnitude Representations Lack Scalar Variability added by Syntology |
2026-04 (from id) |
synthiumjp/weber/m3_pilot/m3_extract.py 6e731249ef09bbf4 |
unverified |
no licence file found · pointer only |
| Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of Lora, Prompt Tuning, and Full Fine-Tuning added by Syntology |
2026-03 (from id) |
eracoding/llm-medical-summarization/src/models/model_loader.py ca3dd3a359eeb3eb |
unverified |
no licence file found · pointer only |
| VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval added by Syntology |
2026-02 (from id) |
dmgcsilva/vigia/inference/vigia/dialogue_sim.py 72ca0249eb720e57 |
unverified |
no licence file found · pointer only |
| Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity added by Syntology |
2026-02 (from id) |
AI21Labs/in-context-ralm/ralm/model_utils.py 93181488645fbb58 |
unverified |
Apache-2.0 (permissive) |
| Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers added by Syntology |
2025-12 (from id) |
cywinski/eliciting-secret-knowledge/sampling/utils.py 7c099df440884000 |
unverified |
MIT (permissive) |
| Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models added by Syntology |
2025-08 (from id) |
ryehr/Consistency/src/tokenization_consistency/models.py e037ac1be5c9f57f |
unverified |
MIT (permissive) |
| LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning |
23 Jun 2025 |
thudm/longwriter/trans_web_demo.py b966ffef888e5413 |
unverified |
Apache-2.0 (permissive) |
| Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking |
18 Feb 2025 |
chuhac/Reasoning-to-Defend/HarmBench/baselines/model_utils.py 090eed81a24d4b32 |
unverified |
MIT (permissive) |
| Adversarial Reasoning at Jailbreaking Time |
3 Feb 2025 |
helloworld10011/adversarial-reasoning/utils.py c46ed68c0b1d6f03 |
unverified |
Apache-2.0 (permissive) |
| LLM Safety Alignment is Divergence Estimation in Disguise |
2 Feb 2025 |
rhaldarpurdue/kldo/eval_utils.py 63fd436cd288e12f |
unverified |
Apache-2.0 (permissive) |
| LLM Safety Alignment is Divergence Estimation in Disguise |
2 Feb 2025 |
rhaldarpurdue/kldo/dataset_generation/compare.py d690eaa700e23094 |
unverified |
Apache-2.0 (permissive) |
| LLM Safety Alignment is Divergence Estimation in Disguise |
2 Feb 2025 |
rhaldarpurdue/kldo/dataset_generation/generate.py b95843089da1efe9 |
unverified |
Apache-2.0 (permissive) |
| Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions |
3 Jan 2025 |
UKPLab/emnlp2025-poate-attack/poate_attack/attacks/utils/models.py 090eed81a24d4b32 |
unverified |
Apache-2.0 (permissive) |
| AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution |
22 Nov 2024 |
r-three/AttriBoT/context_attribution/model_utils.py 4c63c74cb872e812 |
unverified |
no licence file found · pointer only |
| Squeezed Attention: Accelerating Long Context Length LLM Inference |
14 Nov 2024 |
SqueezeAILab/SqueezedAttention/LongBench/pred.py 71856fc92741e6f7 |
unverified |
no licence file found · pointer only |
| Controllable Context Sensitivity and the Knob Behind It |
11 Nov 2024 |
kdu4108/context-vs-prior-finetuning/model_utils/utils.py 63070f500d814722 |
unverified |
no licence file found · pointer only |
| A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement |
17 Oct 2024 |
humainlab/understand_marginpo/analysis/grad_heatmap_1.py 2b221536391e6a55 |
unverified |
no licence file found · pointer only |
| Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks |
5 Oct 2024 |
danshumaan/functional_homotopy/fhautodan/utils/opt_utils.py 576ee553d5cff992 |
ran
|
no licence file found · pointer only |
| Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents |
3 Oct 2024 |
agiresearch/ASB/memory_defense/ppl_utils.py faabedfe90d4f114 |
unverified |
MIT (permissive) |
| A Controlled Study on Long Context Extension and Generalization in LLMs |
18 Sep 2024 |
leooyii/lceg/longbench/pred.py 2b06c68ea29c8f3b |
ran
|
no licence file found · pointer only |
| Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models |
27 Aug 2024 |
Waffle-Liu/DeGCG/baselines/model_utils.py 5949ab432e811362 |
unverified |
MIT (permissive) |
| Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models |
14 Aug 2024 |
chenhuihu/knowledge_in_superposition/Cache_Cov.py 411e60fba7b021a5 |
ran
|
no licence file found · pointer only |
| Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models |
14 Aug 2024 |
chenhuihu/knowledge_in_superposition/Compute_Activation_Cos.py f8698e8cdcb6e36c |
unverified |
no licence file found · pointer only |
| OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection |
23 Jul 2024 |
pku-liang/origen/evaluation/generate_lora.py ab078cc050d0eb8a |
unverified |
no licence file found · pointer only |
| PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment |
23 Jul 2024 |
saltychtao/prealign/src/utils.py c6ea57e16ccfac13 |
unverified |
no licence file found · pointer only |
| Open (Clinical) LLMs are Sensitive to Instruction Phrasings |
12 Jul 2024 |
alceballosa/clin-robust/modules/utils.py 19adfb763485b0b5 |
unverified |
no licence file found · pointer only |
| From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty |
8 Jul 2024 |
mivg/fallbacks/utils/utils.py 03f72536317546c1 |
ran
|
MIT (permissive) |
| ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools |
18 Jun 2024 |
thudm/chatglm/finetune_demo/inference_hf.py 91958a6f1719c4f7 |
unverified |
Apache-2.0 (permissive) |
| ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools |
18 Jun 2024 |
thudm/chatglm/basic_demo/web_demo_gradio.py bb6e915ee912453d |
unverified |
Apache-2.0 (permissive) |
| GUICourse: From General Vision Language Models to Versatile GUI Agents |
17 Jun 2024 |
yiye3/guicourse/Qwen-SFT&Infer/infer.py baaebe9658032c54 |
ran
|
no licence file found · pointer only |
| Probing the Decision Boundaries of In-context Learning in Large Language Models |
17 Jun 2024 |
siyan-zhao/ICL_decision_boundary/get_llm_decision_boundary.py 7d229e134c72cc58 |
unverified |
no licence file found · pointer only |
| Probing the Decision Boundaries of In-context Learning in Large Language Models |
17 Jun 2024 |
siyan-zhao/ICL_decision_boundary/finetune_icl.py fa6cb2e5730316a9 |
unverified |
no licence file found · pointer only |
| Designing a Dashboard for Transparency and Control of Conversational AI |
12 Jun 2024 |
yc015/talktuner-chatbot-llm-dashboard/dashboard_v1/probing/train_probes.py d0a8ad2ed431e1b5 |
unverified |
MIT (permissive) |
| Why Don't Prompt-Based Fairness Metrics Correlate? |
9 Jun 2024 |
chandar-lab/CAIRO/model/model_load.py 58913a7b3b43bf05 |
ran
|
MIT (permissive) |
| Improving Alignment and Robustness with Circuit Breakers |
6 Jun 2024 |
blackswan-ai/circuit-breakers/evaluation/utils.py 2671d81aefc6945e |
unverified |
MIT (permissive) |
| Efficient Adversarial Training in LLMs with Continuous Attacks |
24 May 2024 |
sophie-xhonneux/continuous-advtrain/src/model_utils.py 06cd66993d331abf |
unverified |
MIT (permissive) |
| Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents |
24 Apr 2024 |
kortukov/realistic_knowledge_conflicts/src/model_utils.py b70e091f78d4e3ac |
ran
|
no licence file found · pointer only |
| Massive Activations in Large Language Models |
27 Feb 2024 |
bluorion-com/refine_massive_activations/utils/model.py f9134816cda05113 |
unverified |
MIT (permissive) |
| REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering |
27 Feb 2024 |
RUCAIBox/REAR/rear/src/load_model.py f0ed5a58b6a31117 |
unverified |
no licence file found · pointer only |
| Defending LLMs against Jailbreaking Attacks via Backtranslation |
26 Feb 2024 |
yihanwang617/llm-jailbreaking-defense-backtranslation/attacks/autodan.py 3bcc8e42ea5e457e |
ran
|
BSD-3-Clause (permissive) |
| Coercing LLMs to do and reveal (almost) anything |
21 Feb 2024 |
jonasgeiping/carving/carving/model_interface.py 832c4bb5405dc600 |
unverified |
MIT (permissive) |
| Attacking Large Language Models with Projected Gradient Descent |
14 Feb 2024 |
sigeisler/reinforce-attacks-llms/baselines/model_utils.py 0313a63291b892b7 |
unverified |
MIT (permissive) |
| Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space |
14 Feb 2024 |
schwinnl/llm_embedding_attack/embedding_attack_toxic.py 067aa4bb86ba8911 |
unverified |
MIT (permissive) |
| COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability |
13 Feb 2024 |
Yu-Fangxu/COLD-Attack/opt_util.py 886d1670186fba06 |
ran
|
no licence file found · pointer only |
| PiCO: Peer Review in LLMs based on the Consistency Optimization |
2 Feb 2024 |
PKU-YuanGroup/Peer-review-in-LLMs/llm_judge/common.py 8feea35ea433b832 |
unverified |
no licence file found · pointer only |
| Large Language Models Can Learn Temporal Reasoning |
12 Jan 2024 |
xiongsiheng/tg-llm/src/SFT_TG_Reasoning_ppl.py e9cff349000673db |
unverified |
MIT (permissive) |
| LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples |
2 Oct 2023 |
pku-yuangroup/hallucination-attack/utils.py 8afcadff70823ad7 |
unverified |
MIT (permissive) |
| Graph of Thoughts: Solving Elaborate Problems with Large Language Models |
18 Aug 2023 |
chiral-carbon/kg-for-science/src/utils/utils.py 786ad935afe1d784 |
unverified |
MIT (permissive) |
| OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples |
21 Jul 2023 |
ryuryukke/OUTFOX/utils/utils.py 59ab6a8438c40917 |
unverified |
Apache-2.0 (permissive) |
| A Survey of Large Language Models |
31 Mar 2023 |
xusenlinzy/api-for-open-llm/api/adapter/loader.py 37c3d7ab0a5f9367 |
unverified |
Apache-2.0 (permissive) |
| ContraCLM: Contrastive Learning For Causal Language Model |
3 Oct 2022 |
amazon-science/contraclm/utils.py 0770c863fd59b11f |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.emnlp-main.326 |
|
cl-tohoku/ca-multi-ptn/sft/src/inference.py 6e98cbfbcd92571a |
unverified |
MIT (permissive) |
| arXiv:2023.emnlp-main.771 |
|
Patrick-Ni/KnowEE/src/load_models_and_datasets.py bb16cfebf07908c0 |
unverified |
MIT (permissive) |