| Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology added by Syntology |
2026-05 (from id) |
fdeberna/chat-ptolemaic/evaluation/generate_eval.py 4027d245a8e4b2a5 |
ran
|
no licence file found · pointer only |
| TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models added by Syntology |
2026-03 (from id) |
ZevineXu/TAO-Attack/api_experiments/evaluate_api_models.py 645890ba00019a69 |
unverified |
MIT (permissive) |
| Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model added by Syntology |
2026-02 (from id) |
AndrewWTY/SecCoderX/vul_induce_prompt_pipeline/generate_instructions_batch.py 075ea2984064fe95 |
unverified |
no licence file found · pointer only |
| CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval added by Syntology |
2026-01 (from id) |
idirlab/CaseFacts/experiments/batch_inference_moe.py 824e0abf76e43ed8 |
unverified |
no licence file found · pointer only |
| MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus added by Syntology |
2026-01 (from id) |
yxduir/MCGA/eval/api_mcga.py f5df9a3ab0542ab1 |
unverified |
no licence file found · pointer only |
| Iterative Prompt Refinement for Safer Text-to-Image Generation added by Syntology |
2025-09 (from id) |
ku-dmlab/IPR/vl_rl.py e9c9a97fe7d56a76 |
unverified |
Apache-2.0 (permissive) |
| ImgEdit: A Unified Image Editing Dataset and Benchmark |
26 May 2025 |
pku-yuangroup/imgedit/Benchmark/Basic/basic_bench.py 913a857e957a8c91 |
ran · our draft was wrong
|
no licence file found · pointer only |
| DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution |
22 May 2025 |
zhengchen1999/DOVE/finetune/datasets/utils.py fad222ebf1926ea1 |
ran
|
Apache-2.0 (permissive) |
| Better Estimation of the KL Divergence Between Language Models |
14 Apr 2025 |
rycolab/kl-rb/rloo_sentiment.py 8c77fc1b61189419 |
ran · our draft was wrong
|
no licence file found · pointer only |
| RewardSDS: Aligning Score Distillation via Reward-Weighted Sampling |
12 Mar 2025 |
itaychachy/RewardSDS/evaluation/aesthetic_eval.py d98daaa9e9aeb095 |
unverified |
MIT (permissive) |
| Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think |
2 Mar 2025 |
Chuge0335/EDG/evaluation/infer_multi.py ea66ba28a2e932e1 |
ran · our draft was wrong
|
MIT (permissive) |
| LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models |
10 Feb 2025 |
jadohu/lantern/entrypoints/generate_images.py 865916f0c52143f1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models |
27 Dec 2024 |
wutao-cs/videomaker/inference.py ea66ba28a2e932e1 |
ran · our draft was wrong
|
licence not identified · pointer only |
| RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework |
24 Oct 2024 |
ewanwong/rsa-control/toxicity_bias_mitigation/io_utils.py 44bbf4459b4f2e94 |
unverified |
MIT (permissive) |
| Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino |
20 Sep 2024 |
aisingapore/kalahi/kalahi/utilities.py a018415e35b1d137 |
ran
|
CC-BY-4.0 · pointer only |
| Score Forgetting Distillation: A Swift, Data-Free Method for Machine Unlearning in Diffusion Models |
17 Sep 2024 |
ml-research/i2p/eval/q16.py 19896f44b0b6ff3e |
unverified |
no licence file found · pointer only |
| MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector |
16 Aug 2024 |
wjfu99/mia-tuner/my_utils.py 91d1d9b0c9c266d7 |
ran
|
no licence file found · pointer only |
| CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer |
12 Aug 2024 |
thudm/cogvideo/finetune/datasets/utils.py fad222ebf1926ea1 |
ran
|
Apache-2.0 (permissive) |
| BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models |
21 Jul 2024 |
bigbench2024/bigbench2024/benchmark/generate/generate.py a2205613ccd243e1 |
ran
|
GPL-3.0 (copyleft) · pointer only |
| Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training |
12 Jul 2024 |
robustnlp/derta/evaluation.py ce51d685abf1ad76 |
ran · our draft was wrong
|
MIT (permissive) |
| Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses |
3 Jun 2024 |
sail-sg/I-FSJ/api_experiments/evaluate_api_models.py 645890ba00019a69 |
unverified |
MIT (permissive) |
| HonestLLM: Toward an Honest and Helpful Large Language Model |
1 Jun 2024 |
Flossiee/HonestyLLM/training_free/prompt_loader.py c80548e335885a3e |
ran
|
no licence file found · pointer only |
| FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models |
28 May 2024 |
astarojth/faintbench-v1/comfyui/generate_lcm.py 632c123d9b1dc4f8 |
ran
|
GPL-3.0 (copyleft) · pointer only |
| Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks |
17 May 2024 |
C-anwoy/Cross-Task-ICL/utils/data.py e401d40128e5c42a |
ran
|
MIT (permissive) |
| UniRAG: Universal Retrieval Augmentation for Large Vision Language Models |
16 May 2024 |
castorini/unirag/src/unirag/eval_image_generation.py 1d3391d39efd8b4c |
ran
|
no licence file found · pointer only |
| MedAdapter: Efficient Test-Time Adaptation of Large Language Models towards Medical Reasoning |
5 May 2024 |
wshi83/MedAdapter/reward_model/orm/orm_guide.py 537590c07503fa34 |
unverified |
no licence file found · pointer only |
| Protecting Your LLMs with Information Bottleneck |
22 Apr 2024 |
llm-attacks/llm-attacks/api_experiments/evaluate_api_models.py 645890ba00019a69 |
unverified |
MIT (permissive) |
| StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows |
17 Mar 2024 |
kevin666aa/stateflow/ALFWorld/src/chat_utils.py ad7277e0c33fc9aa |
ran
|
no licence file found · pointer only |
| Defending LLMs against Jailbreaking Attacks via Backtranslation |
26 Feb 2024 |
yihanwang617/llm-jailbreaking-defense-backtranslation/utils.py 37cddda277e8045d |
ran
|
BSD-3-Clause (permissive) |
| TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification |
20 Feb 2024 |
framartin/trap/detect_llm/baseline_ppl.py c7dc4896ac741214 |
ran
|
MIT (permissive) |
| TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification |
20 Feb 2024 |
parameterlab/trap/llm_attacks/api_experiments/evaluate_api_models.py 645890ba00019a69 |
unverified |
MIT (permissive) |
| You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments |
16 Nov 2023 |
orange0629/llm-personas/src/prompt-variation/run_prompts_on_models.py 872da3bc878987a3 |
ran
|
no licence file found · pointer only |
| BLESS: Benchmarking Large Language Models on Sentence Simplification |
24 Oct 2023 |
ZurichNLP/BLESS/utils/helpers.py d2360f86ab2cb2b0 |
ran
|
MIT (permissive) |
| Three Bricks to Consolidate Watermarks for Large Language Models |
26 Jul 2023 |
facebookresearch/three_bricks/main_watermark.py e43d8e6c09e3a48a |
ran · our draft was wrong
|
licence not identified · pointer only |
| Towards Safe Self-Distillation of Internet-Scale Text-to-Image Diffusion Models |
12 Jul 2023 |
nannullna/safe-diffusion/generate.py 9f927d43843c43af |
ran · our draft was wrong
|
MIT (permissive) |
| Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning |
7 May 2023 |
sf-zhai/badt2i/evaluation/img_gen.py bbd2ac7ff643d1cf |
unverified |
Apache-2.0 (permissive) |
| ReCode: Robustness Evaluation of Code Generation Models |
20 Dec 2022 |
frabbisw/robustextended/evalplus/evaluate_inputs.py d459632bb421c25b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Challenges in Measuring Bias via Open-Ended Language Generation |
23 May 2022 |
feyzaakyurek/bias-textgen/complete_prompts.py 75208746d14a8b86 |
unverified |
MIT (permissive) |
| Quantifying Memorization Across Neural Language Models |
15 Feb 2022 |
ftramer/lm-extraction-benchmark/baseline/simple_baseline.py 53e3e08755cfa895 |
unverified |
Apache-2.0 (permissive) |
| MLPerf Inference Benchmark |
6 Nov 2019 |
mlcommons/inference/text_to_video/wan-2.2-t2v-a14b/run_inference.py 7b3e0f2610f104fb |
unverified |
Apache-2.0 (permissive) |
| arXiv:openreview_47JZSOkw5C |
|
bxuanz/SigMa/run_sigma_qwen.py 63ccd86a50d62f8e |
unverified |
MIT (permissive) |
| arXiv:2025.naacl-long.14 |
|
vectara/mirage-bench/mirage_bench/util.py ee845a8455f6d910 |
unverified |
Apache-2.0 (permissive) |