| The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction added by Syntology |
2026-09 (from id) |
hopit-ai/Moda_ner/suite/catalog/score.py fa9634889181f1b8 |
unverified |
MIT (permissive) |
| Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning added by Syntology |
2026-09 (from id) |
kaist-dmlab/CluSTER/src/CluSTER/utils.py 764fe981599f1c88 |
unverified |
Apache-2.0 (permissive) |
| T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation added by Syntology |
2026-09 (from id) |
LLMSecResearch/T2LSC-Bench/evaluation/metrics.py 64690e0ae399636c |
unverified |
no licence file found · pointer only |
| EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision added by Syntology |
2026-09 (from id) |
18277390221/EmoStance/src/latent_stance_control/evaluate_text_quality.py 33d006261b75ac74 |
unverified |
MIT (permissive) |
| StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions added by Syntology |
2026-09 (from id) |
Cha0Ga0/SWAPSTATE/src/hs_swap/io.py 725cbdee82cfaf24 |
unverified |
no licence file found · pointer only |
| RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation added by Syntology |
2026-09 (from id) |
ZhongruChen/RPCBench/src/step1/_step1_common.py 569586b8eea9a84e |
unverified |
no licence file found · pointer only |
| RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation added by Syntology |
2026-09 (from id) |
ZhongruChen/RPCBench/src/step2_runner/_runner_common.py c52ed7dd243c3812 |
unverified |
no licence file found · pointer only |
| The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems added by Syntology |
2026-09 (from id) |
mpi-dsg/irreversibility-budget/sim/traces/plot.py 953febbb771fc0e8 |
unverified |
MIT (permissive) |
| Language Chain in Alignment: Cross-lingual Ranking Preference Optimization added by Syntology |
2026-08 (from id) |
dltmddbs100/CRPO/dataset/prepare_dataset.py 260bb4a4a14529db |
ran
|
MIT (permissive) |
| KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search added by Syntology |
2026-08 (from id) |
back-kh/Khmer-Semantic-Search/codes/evaluate_retrieval.py 4fc027c853affddb |
ran · our draft was wrong
|
MIT (permissive) |
| EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering added by Syntology |
2026-08 (from id) |
RamonMeng/EnSI-RAG/ensi-rag-financebench-eval/ensi_financebench/qa_engine.py c24e644ee0c0b239 |
ran
|
no licence file found · pointer only |
| RepBench: A Benchmark for Representation Engineering added by Syntology |
2026-07 (from id) |
xlands/RepBench/src/representation_cluster/representation_fit_pipeline.py 7902fbf13b57a305 |
ran
|
Apache-2.0 (permissive) |
| Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models added by Syntology |
2026-07 (from id) |
js-lee-AI/answer-leakage/answer_leakage/corpus.py 529acb6cf747d1fc |
ran
|
MIT (permissive) |
| Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? added by Syntology |
2026-06 (from id) |
THU-KEG/RuVerBench/code/run_judges/agenticcoding_runner.py 704044cedacb5584 |
ran
|
licence not identified · pointer only |
| Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? added by Syntology |
2026-06 (from id) |
THU-KEG/RuVerBench/code/validate_package.py e7db7ab88e7a0bd2 |
unverified |
licence not identified · pointer only |
| How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring added by Syntology |
2026-06 (from id) |
mathieu0905/Code-Anchor/evaluation/eval_localization_codex.py 14b1625bd8d3d013 |
unverified |
Apache-2.0 (permissive) |
| From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents added by Syntology |
2026-06 (from id) |
AI-Santiago/KAware/run_function_call_minimal.py 4afd53f9866aae14 |
ran
|
no licence file found · pointer only |
| Assessing Reliability of Symbol Detection in Concept Bottleneck Models added by Syntology |
2026-06 (from id) |
Fuminides/cbm_sanity/src/symbol_sanity/multihead.py 5b2dd5de02dc3211 |
ran · our draft was wrong
|
MIT (permissive) |
| Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? added by Syntology |
2026-06 (from id) |
CHATS-lab/coding-agent-safety-monitor/utils/io.py 8add8881bedf4a28 |
ran
|
MIT (permissive) |
| Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling added by Syntology |
2026-06 (from id) |
kaist-cvml/perception-judge/prepare-datasets/post_processing_merge_folders.py 2d6a14cf6379542d |
unverified |
no licence file found · pointer only |
| MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding added by Syntology |
2026-05 (from id) |
xiaofengShi/MechVQA/data_generation/generate_vqa_free_query.py f73db1bba0fc60bf |
ran
|
Apache-2.0 (permissive) |
| MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding added by Syntology |
2026-05 (from id) |
xiaofengShi/MechVQA/data_generation/check_generated_questions.py 0d8d1a6b7845a85d |
unverified |
Apache-2.0 (permissive) |
| Analyzing Quality-Latency-Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation added by Syntology |
2026-05 (from id) |
EugPal/rag-lora-tradeoffs/src/data_pipeline/build_kubernetes_qa_dataset.py 65440b243bca4da1 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning added by Syntology |
2026-05 (from id) |
AlexFanw/LegalSearch-R1/user/calculate_metrics.py 27498a27e49ed413 |
ran
|
Apache-2.0 (permissive) |
| Multi-Rollout On-Policy Distillation via Peer Successes and Failures added by Syntology |
2026-05 (from id) |
viviable/mopd_code/analysis/score_teacher_contexts.py 903f3f3a478c7f63 |
ran
|
Apache-2.0 (permissive) |
| Multi-Rollout On-Policy Distillation via Peer Successes and Failures added by Syntology |
2026-05 (from id) |
viviable/mopd_code/analysis/compute_teacher_signal_metrics.py 2b6d83de04b5e4ad |
ran
|
Apache-2.0 (permissive) |
| Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration added by Syntology |
2026-05 (from id) |
PKU-SEC-Lab/SPEX/bfs_async.py 951d55a786237f31 |
ran · our draft was wrong
|
no licence file found · pointer only |
| OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories added by Syntology |
2026-05 (from id) |
PolarSeeker/OpenSeeker/eval/generate_answer.py 01c02f8e0d457fb1 |
ran
|
no licence file found · pointer only |
| Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning added by Syntology |
2026-04 (from id) |
DJC-GO-SOLO/Latent-SFT/eval/eval_latent_model_hf_batch.py 3d5343ae289698e8 |
ran · our draft was wrong
|
MIT (permissive) |
| On the Step Length Confounding in LLM Reasoning Data Selection added by Syntology |
2026-04 (from id) |
wangbing1416/ASLEC/merge_cal_limo.py 0a2bd131350902b8 |
unverified |
no licence file found · pointer only |
| On the Step Length Confounding in LLM Reasoning Data Selection added by Syntology |
2026-04 (from id) |
wangbing1416/ASLEC/select_sft_data.py 8de9a2e11d7d03fa |
unverified |
no licence file found · pointer only |
| On the Step Length Confounding in LLM Reasoning Data Selection added by Syntology |
2026-04 (from id) |
wangbing1416/ASLEC/select_sft_data_ours.py 8fc83dfd8da40551 |
unverified |
no licence file found · pointer only |
| Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy added by Syntology |
2026-04 (from id) |
ku262/EndoASR/eval/eval_acc.py 41cce0cc5f4d5c38 |
unverified |
no licence file found · pointer only |
| Meta-Reinforcement Learning with Self-Reflection for Agentic Search added by Syntology |
2026-03 (from id) |
tengxiao1/MR-Search/meta-search/search/retrieval_server.py 02171abe8492fcfd |
ran · our draft was wrong
|
no licence file found · pointer only |
| Meta-Reinforcement Learning with Self-Reflection for Agentic Search added by Syntology |
2026-03 (from id) |
tengxiao1/MR-Search/meta-search/search/retrieval.py e50ece300cf79a17 |
unverified |
no licence file found · pointer only |
| Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges added by Syntology |
2026-02 (from id) |
ZDCSlab/Rubrics-as-an-Attack-Surface/downstream_eval/eval/scores.py d97bc4cdc990894c |
unverified |
MIT (permissive) |
| Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges added by Syntology |
2026-02 (from id) |
ZDCSlab/Rubrics-as-an-Attack-Surface/downstream_eval/eval/analyze.py ffbf5293f2483abe |
unverified |
MIT (permissive) |
| Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges added by Syntology |
2026-02 (from id) |
ZDCSlab/Rubrics-as-an-Attack-Surface/downstream_eval/eval/generate.py f07107c7894990c1 |
unverified |
MIT (permissive) |
| Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges added by Syntology |
2026-02 (from id) |
ZDCSlab/Rubrics-as-an-Attack-Surface/downstream_eval/eval/select_best.py 6a94611fb33432fd |
unverified |
MIT (permissive) |
| REVIS: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models added by Syntology |
2026-02 (from id) |
antgroup/Revis/utils/chair.py 098ce1dbb8a719bf |
unverified |
Apache-2.0 (permissive) |
| ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents added by Syntology |
2026-02 (from id) |
pc-inno/ICA_MM_deepsearch/llm_judge.py f45882c6806cb8d4 |
unverified |
no licence file found · pointer only |
| Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference added by Syntology |
2026-02 (from id) |
Lancelot-Xie/DRIFT/src/drift/inference/eval_multi.py c35fa39d6741b7b7 |
unverified |
no licence file found · pointer only |
| Alignment-Aware Model Adaptation via Feedback-Guided Optimization added by Syntology |
2026-02 (from id) |
facebookresearch/TruthRL/data_utils/retrieval.py e50ece300cf79a17 |
unverified |
no licence file found · pointer only |
| On the Paradoxical Interference between Instruction-Following and Task Solving added by Syntology |
2026-01 (from id) |
kijlk/IF-Interference/src/math_and_qa/evaluation.py 9190ce603f740a88 |
unverified |
no licence file found · pointer only |
| Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented Generation added by Syntology |
2026-01 (from id) |
Cherry-qwq/LcRL-Open/search_r1/search/retrieval.py e50ece300cf79a17 |
unverified |
Apache-2.0 (permissive) |
| GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization added by Syntology |
2026-01 (from id) |
zzy1127/GIFT/utils/data_utils.py 82f198e67affc0bb |
unverified |
no licence file found · pointer only |
| Towards Understanding Valuable Preference Data for Large Language Model Alignment added by Syntology |
2025-10 (from id) |
tmlr-group/TIF_LossDiff-IRM/winrate_eval/single_score.py a65ff6fa4d9d7cc7 |
unverified |
no licence file found · pointer only |
| Towards Understanding Valuable Preference Data for Large Language Model Alignment added by Syntology |
2025-10 (from id) |
tmlr-group/TIF_LossDiff-IRM/winrate_eval/single_score_local.py 07adb23040d2275c |
unverified |
no licence file found · pointer only |
| Towards Understanding Valuable Preference Data for Large Language Model Alignment added by Syntology |
2025-10 (from id) |
tmlr-group/TIF_LossDiff-IRM/analysis/data_select_by2_mid_k.py 17f3dfa5c8d12d87 |
unverified |
no licence file found · pointer only |
| Towards Understanding Valuable Preference Data for Large Language Model Alignment added by Syntology |
2025-10 (from id) |
tmlr-group/TIF_LossDiff-IRM/analysis/data_select_mid_k.py e2a67da3a04cd3c7 |
unverified |
no licence file found · pointer only |
| Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models added by Syntology |
29 Sep 2025 |
scxue/advantage_weighted_matching/advantage_weighted_matching/dataset/merge_genevaltask.py 90a22cdf7d22703b |
unverified |
Apache-2.0 (permissive) |
| ConfTuner: Training Large Language Models to Express Their Confidence Verbally added by Syntology |
2025-08 (from id) |
liushiliushi/ConfTuner/src/llama_recipes/datasets2/gsm8k_dataset.py 951d55a786237f31 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction added by Syntology |
2025-08 (from id) |
Xinhe-Li/LoRID/src/utils/file_utils.py 19fb595831e42a78 |
unverified |
Apache-2.0 (permissive) |
| WebSailor: Navigating Super-human Reasoning for Web Agent |
3 Jul 2025 |
alibaba-nlp/webagent/WebAgent/NestBrowse/utils.py ecb8ee265d233ddd |
unverified |
Apache-2.0 (permissive) |
| TaskCraft: Automated Generation of Agentic Tasks |
11 Jun 2025 |
oppo-personalai/taskcraft/taskcraft/utils.py c5bf8e54f0728275 |
unverified |
MIT (permissive) |
| AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting |
24 May 2025 |
joeying1019/adactrl/utils/data_utils.py 82f198e67affc0bb |
unverified |
Apache-2.0 (permissive) |
| Distilling LLM Agent into Small Models with Retrieval and Code Tools |
23 May 2025 |
Nardien/agent-distillation/search/retriever_server.py 02171abe8492fcfd |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs |
21 May 2025 |
mira-ai-lab/deliberation-on-priors/reasoning/instantiation.py 3d5343ae289698e8 |
ran · our draft was wrong
|
MIT (permissive) |
| Dynamic Early Exit in Reasoning Models |
22 Apr 2025 |
iie-ycx/deer/vllm-deer.py c6485c4cde6c2322 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning |
10 Apr 2025 |
tiger-ai-lab/vl-rethinker/openrlhf/trainer/evaluator.py 059f7ea3562486c2 |
unverified |
Apache-2.0 (permissive) |
| LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQL |
24 Mar 2025 |
Satissss/LinkAlign/preprocess.py a22422413f1c610b |
ran
|
MIT (permissive) |
| Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings |
19 Mar 2025 |
salesforceairesearch/contextualjudgebench/utils/utils.py c89138f7dd4645e9 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence |
19 Feb 2025 |
lux0926/asprm/evaluation/code/TVD/tvd_lcb.py 62569752807c6bc6 |
ran · our draft was wrong
|
no licence file found · pointer only |
| How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training |
16 Feb 2025 |
zjunlp/dynamicknowledgecircuits/utils.py bf43819905001d75 |
unverified |
MIT (permissive) |
| CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging |
8 Feb 2025 |
kagnlp/CodeGenerator/src/datasets/convert-apps-xcode.py 94c65e47ba360f89 |
unverified |
MIT (permissive) |
| VideoRoPE: What Makes for Good Video Rotary Position Embedding? |
7 Feb 2025 |
wiselnn570/videorope/eval/model_longvideobench_qwen2_vl.py 15f446d6c10dc173 |
unverified |
Apache-2.0 (permissive) |
| A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing |
2024-12 (from id) |
iSEngLab/LLM4UT_Empirical/Inference_Script/open_source.py 02171abe8492fcfd |
ran · our draft was wrong
|
no licence file found · pointer only |
| UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models |
16 Dec 2024 |
amourwaltz/ualign/code/utils.py d5792c4c85064ef5 |
unverified |
no licence file found · pointer only |
| Classifier-free guidance in LLMs Safety |
8 Dec 2024 |
rgsmirnov/cfg_safety_llm/train_orpo.py 435e1bfd0feee04a |
ran · our draft was wrong
|
no licence file found · pointer only |
| DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling |
6 Dec 2024 |
mozerwang/demo/src/utils/functions.py aa2475b60b65226f |
unverified |
Apache-2.0 (permissive) |
| Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code |
3 Dec 2024 |
jetbrains-research/pandasplotbench/plotting_benchmark/vis_generator.py e661693417b2d5b1 |
unverified |
Apache-2.0 (permissive) |
| CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking |
1 Dec 2024 |
gangiswag/cornstack/src/evaluations/eval_localization.py a4312a4596349487 |
unverified |
Apache-2.0 (permissive) |
| FullStack Bench: Evaluating LLMs as Full Stack Coders |
30 Nov 2024 |
bytedance/fullstackbench/src/utils.py ff3dacdd2cf428ba |
unverified |
Apache-2.0 (permissive) |
| Star Attention: Efficient LLM Inference over Long Sequences |
26 Nov 2024 |
NVIDIA/Star-Attention/run_star_attn_inference.py 2a87c292d08948bc |
unverified |
Apache-2.0 (permissive) |
| SelfCodeAlign: Self-Alignment for Code Generation |
31 Oct 2024 |
bigcode-project/starcoder2-self-align/src/star_align/utils.py b0fdb2b203090d2e |
unverified |
Apache-2.0 (permissive) |
| ALTA: Compiler-Based Analysis of Transformers |
23 Oct 2024 |
google-deepmind/alta/framework/common/io_utils.py bafd744639636cc4 |
unverified |
Apache-2.0 (permissive) |
| IPO: Interpretable Prompt Optimization for Vision-Language Models |
20 Oct 2024 |
lmsdss/ipo/optimization/eval_utils.py a4cf9d0d1ece601e |
ran · our draft was wrong
|
no licence file found · pointer only |
| LoGU: Long-form Generation with Uncertainty Expressions |
18 Oct 2024 |
rhyang2021/logu/utils.py a05eb9a0d2147891 |
ran
|
no licence file found · pointer only |
| RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards |
17 Oct 2024 |
openmatch/rag-ddr/src/generator/gen_dpo_data.py a22422413f1c610b |
ran
|
MIT (permissive) |
| RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards |
17 Oct 2024 |
openmatch/rag-ddr/src/generator/gen_llm_response.py fdd1249aa05ba5f3 |
ran
|
MIT (permissive) |
| RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards |
17 Oct 2024 |
openmatch/rag-ddr/src/knowledgeRefinement/kr_inference.py 90a22cdf7d22703b |
unverified |
MIT (permissive) |
| JudgeBench: A Benchmark for Evaluating LLM-based Judges |
16 Oct 2024 |
ScalerLab/JudgeBench/utils/file_operations.py 195bbcd43396ba20 |
ran
|
no licence file found · pointer only |
| How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective |
14 Oct 2024 |
tengxiao1/GSIL/gsil/combine.py ed3cbe8d977dbf1d |
ran
|
no licence file found · pointer only |
| Toward General Instruction-Following Alignment for Retrieval-Augmented Generation |
12 Oct 2024 |
dongguanting/FollowRAG/FollowRAG/utils/util.py 7bb90ae0136ac9dd |
ran
|
no licence file found · pointer only |
| Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning |
10 Oct 2024 |
shuhewang1998/packing-analysis/transfer_data.py acd6bebe90eafbda |
ran
|
no licence file found · pointer only |
| Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought |
8 Oct 2024 |
LightChen233/reasoning-boundary/utils/mm_tool.py 8f0ae0bd1b3d8a66 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought |
8 Oct 2024 |
lightchen233/reasoning-boundary/draw_bound_text.py 86262555b2a8871d |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models |
3 Oct 2024 |
zhipeixu/fakeshield/MFLM/cli_demo.py b1c0bb4ee3a78693 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models |
2 Oct 2024 |
tony10101105/ExpEmergence/evaluation/abstract_narrative_understanding/abstract_narrative_understanding_question_grouping.py 69f3426a987055c9 |
ran
|
MIT (permissive) |
| EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models |
26 Sep 2024 |
MaLA-LM/emma-500/evaluation/Aya/evaluate.py 47ad7e8829ffe863 |
ran
|
no licence file found · pointer only |
| IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS |
9 Sep 2024 |
ai4bharat/indicvoices-r/VoiceCraft/inference.py c0e30f6224e1afda |
ran
|
CC-BY-4.0 · pointer only |
| SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding |
28 Aug 2024 |
dptech-corp/Uni-SMART/SciLitLLM/cpt/quality_control/llama_infer.py 793cd3b000794433 |
ran
|
MIT (permissive) |
| Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack |
23 Jul 2024 |
ink-usc/lifelong-icl/dataloader/incremental_task.py 9cfdd3d5a62805fa |
ran · our draft was wrong
|
MIT (permissive) |
| Case2Code: Learning Inductive Reasoning with Synthetic Data |
17 Jul 2024 |
choosewhatulike/case2code/api_call_util.py 65c57366104cbd60 |
ran
|
Apache-2.0 (permissive) |
| Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression |
6 Jul 2024 |
zhichaoxu-shufe/beyond-perplexity-compression-safety-eval/src/dataset.py 39f2050e41061d06 |
ran
|
no licence file found · pointer only |
| RouteLLM: Learning to Route LLMs with Preference Data |
26 Jun 2024 |
lm-sys/routellm/routellm/evals/gsm8k/generate_responses.py d6397718adba76b0 |
ran
|
Apache-2.0 (permissive) |
| Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track |
24 Jun 2024 |
castorini/nuggetizer/src/nuggetizer/cli/io.py 719779e74d1ac660 |
unverified |
Apache-2.0 (permissive) |
| Multimodal Table Understanding |
12 Jun 2024 |
spursgozmy/table-llava/llava/eval/generate_webpage_data_from_table.py 7beeb01acc36919e |
ran
|
Apache-2.0 (permissive) |
| An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection |
10 Jun 2024 |
datasec-lab/codebreaker/CodeGen/codegen1/benchmark/mtpb_exec.py b196c82bf4ad6d4d |
ran · our draft was wrong
|
no licence file found · pointer only |
| TESTEVAL: Benchmarking Large Language Models for Test Case Generation |
2024-06 (from id) |
llm4softwaretesting/testeval/data_utils.py 60c584f55c9451fc |
ran
|
MIT (permissive) |
| X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions |
30 May 2024 |
znlp/x-instruction/utils/chatgpt_generation.py 6356690e9f0d26f0 |
ran
|
MIT (permissive) |
| PromptWizard: Task-Aware Prompt Optimization Framework |
28 May 2024 |
microsoft/promptwizard/promptwizard/glue/paramlogger/file_utils.py a18ff17da97103cd |
unverified |
MIT (permissive) |
| AutoPSV: Automated Process-Supervised Verifier |
27 May 2024 |
rookie-joe/autocv/utils/verfier_datasets.py c409d17fa33773fa |
ran
|
no licence file found · pointer only |
| M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought |
26 May 2024 |
LightChen233/M3CoT/utils/data.py 8f0ae0bd1b3d8a66 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Intruding with Words: Towards Understanding Graph Injection Attacks at the Text Level |
26 May 2024 |
Leirunlin/Text-level-Graph-Attack/LLM_utils.py f6dadc26879ae08a |
ran
|
no licence file found · pointer only |
| VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding |
22 May 2024 |
gyxxyg/vtg-llm/utils/get_coco_format.py 8871a34f67963cf7 |
unverified |
Apache-2.0 (permissive) |
| xFinder: Robust and Pinpoint Answer Extraction for Large Language Models |
20 May 2024 |
IAAR-Shanghai/xFinder/scripts/dataset_construction/create_benchmark_dataset.py 352121e0de5ba153 |
ran · our draft was wrong
|
licence not identified · pointer only |
| Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards |
16 Apr 2024 |
hbin0701/Self-Explore/gen/get_rft_data.py c4783f2528e590ab |
ran
|
no licence file found · pointer only |
| Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer |
5 Apr 2024 |
tartunlp/llammas/scripts/chat_data/instructions_to_chat.py 9d32f043536de9b2 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided |
1 Apr 2024 |
honglizhan/resort_cognitive_reappraisal/utils.py ca3d4025d7524b34 |
ran
|
no licence file found · pointer only |
| Beyond Embeddings: The Promise of Visual Table in Visual Reasoning |
27 Mar 2024 |
lavi-lab/visual-table/llava/eval/generate_webpage_data_from_table.py 7beeb01acc36919e |
ran
|
Apache-2.0 (permissive) |
| Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization |
26 Mar 2024 |
jinpz/dtv/formalize_proof.py 1eee05718afec3bf |
ran · our draft was wrong
|
MIT (permissive) |
| LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models |
20 Mar 2024 |
Rcrossmeister/RLQG/model/converter.py 2a7c632d17841b02 |
unverified |
MIT (permissive) |
| Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates |
28 Feb 2024 |
vfleaking/PTST/gpt-api/data_utils/prep_data.py 64c1b0186895ea4a |
ran
|
Apache-2.0 (permissive) |
| Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models |
21 Feb 2024 |
zwhe99/x-sir/utils.py 8f17a72bd2c9a658 |
ran
|
no licence file found · pointer only |
| Towards Building Multilingual Language Model for Medicine |
21 Feb 2024 |
magic-ai4med/mmedlm/inference/inference.py 7aa433b5d2a47cd0 |
ran · our draft was wrong
|
licence not identified · pointer only |