| FrameBench:A Language Understanding Benchmark Based on Frame Semantics added by Syntology |
2026-09 (from id) |
SasanoLab/FrameBench/src/generation/utils/utils.py 22a081a82f22f469 |
unverified |
no licence file found · pointer only |
| Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories added by Syntology |
2026-09 (from id) |
nabirarashid/structural-retrieval/src/data.py 88655287e149326b |
unverified |
licence not identified · pointer only |
| Reactivating Test-Time Scaling for Plane Geometry Problem Solving added by Syntology |
2026-08 (from id) |
Jason8Kang/ReTTS-PGPS/src/retts_pgp/eval/evaluate.py 5782d82a3417e0a7 |
unverified |
licence not identified · pointer only |
| Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning added by Syntology |
2026-08 (from id) |
SolereZhang/GC-OPD/evaluation/aggregate_main_table.py 5627227cf9041a8d |
ran
|
Apache-2.0 (permissive) |
| MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques added by Syntology |
2026-08 (from id) |
WuqnEl/MuseCritic/eval/compute_corr.py 90450e4180be9ae8 |
ran
|
no licence file found · pointer only |
| Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness added by Syntology |
2026-08 (from id) |
rsrijith/cikm26-catalog-faithfulness/code/prep_yelp.py 7a1382f2ea061f72 |
ran
|
licence not identified · pointer only |
| OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories added by Syntology |
2026-08 (from id) |
Changhao-Xiang/OpenVisTool/distill/filter/evaluate_instructive_trajectory_ablation.py fd449ff795ee3154 |
ran
|
MIT (permissive) |
| OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories added by Syntology |
2026-08 (from id) |
Changhao-Xiang/OpenVisTool/distill/filter/filter_tool_gain_with_prefix.py fa52ba8cf457edcf |
ran
|
MIT (permissive) |
| When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO added by Syntology |
2026-08 (from id) |
CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO/analyze_mechanism.py f835ffe0b6606634 |
ran
|
MIT (permissive) |
| When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO added by Syntology |
2026-08 (from id) |
CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO/eval_full_clusters.py 40d46e8d30cf8006 |
ran
|
MIT (permissive) |
| When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO added by Syntology |
2026-08 (from id) |
CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO/eval_vllm.py 70091d81e0d906bd |
ran
|
MIT (permissive) |
| SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay added by Syntology |
2026-08 (from id) |
125jx/SeqLLM/evaluation/eval_item_understanding_task/infer_recprobe_vllm.py ffd4c6463dd0a477 |
ran
|
no licence file found · pointer only |
| HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning added by Syntology |
2026-08 (from id) |
JiananXie/HopRefusalBench/evaluate.py 2820cd4e6747f653 |
ran
|
no licence file found · pointer only |
| Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation added by Syntology |
2026-07 (from id) |
Wuzheng02/ESPP/pqa_construction/common.py 0b1828f7666144a8 |
ran
|
no licence file found · pointer only |
| Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation added by Syntology |
2026-07 (from id) |
Wuzheng02/ESPP/scoring_pipeline/common.py ecfe93b1c22b03a1 |
ran
|
no licence file found · pointer only |
| Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation added by Syntology |
2026-07 (from id) |
eth-sri/rlm-training-merging/src/evaluation/text_summarization.py 455c557e67fba06c |
ran
|
no licence file found · pointer only |
| A Unified Benchmarking Framework for Spatiotemporal Event Modeling added by Syntology |
2026-07 (from id) |
YahyaAalaila/seahorse/seahorse/utils.py fd2e5730d6853aae |
ran
|
Apache-2.0 (permissive) |
| Geometric Measurements of the Axiom of Choice in Neural Proof Embeddings added by Syntology |
2026-06 (from id) |
rodrgo/geometric-axiom-of-choice/experiments/classical_ablation_analyze.py 83d5da6867b97795 |
ran
|
no licence file found · pointer only |
| How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring added by Syntology |
2026-06 (from id) |
mathieu0905/Code-Anchor/evaluation/eval_metric.py 846475125bc11ceb |
ran
|
Apache-2.0 (permissive) |
| Localizing RL-Induced Tool Use to a Single Crosscoder Feature added by Syntology |
2026-06 (from id) |
Antebe/model_diffing_crosscoders/analysis/run_xcoder_hparams_analysis.py b7bc84fc9ab32ca0 |
ran
|
no licence file found · pointer only |
| Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods added by Syntology |
2026-06 (from id) |
francescacarlon/Thinking-Like-a-Scientist/src/pipeline/compare_model_swap.py 041e02ba2162b2a0 |
ran
|
MIT (permissive) |
| Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation added by Syntology |
2026-06 (from id) |
PunchlineAAAA/DICE/ReasonAug/bright_verify.py 58e81c79c7830cc9 |
ran
|
MIT (permissive) |
| Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes added by Syntology |
2026-06 (from id) |
MohamedBayan/MemeReason/annotation/azure_batch.py bcc2eb61ab57d401 |
ran
|
no licence file found · pointer only |
| A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets added by Syntology |
2026-06 (from id) |
Tej-55/NAPE/finetuning/data_preparation.py cf97dcd92963847e |
ran
|
licence not identified · pointer only |
| How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation added by Syntology |
2026-06 (from id) |
fensorechase/rag-diverse-benchmarks-synthetic-qa/rag_system/generate/complete_analysis.py e5a2f31357586772 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output added by Syntology |
2026-06 (from id) |
lmarena/arena-hard-auto/qa_browser.py d049c92b5f5adfa3 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| CARE: A Conformal Safety Layer for Medical Summarization added by Syntology |
2026-06 (from id) |
som-shahlab/CARE/care/utils.py d062537281fbcda4 |
ran
|
no licence file found · pointer only |
| Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So? added by Syntology |
2026-06 (from id) |
LSU-ATHENA/HPM-Predict/gen_dataset/run_hunyuan_rank100_extension.py cfd35648dca1b0b7 |
ran
|
no licence file found · pointer only |
| Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill added by Syntology |
2026-06 (from id) |
Qwen-Applications/Skill-RM/experiments/if_rewardbench/io_utils.py 6f16fa558d7424a4 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| When Model Merging Breaks Routing: Training-Free Calibration for MoE added by Syntology |
2026-06 (from id) |
huangcb01/HARC/src/utils.py c3dd1094f3b76096 |
ran
|
no licence file found · pointer only |
| When Model Merging Breaks Routing: Training-Free Calibration for MoE added by Syntology |
2026-06 (from id) |
huangcb01/HARC/src/merge_method/regmean.py 0d2c2dfbee2f518c |
ran
|
no licence file found · pointer only |
| MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding added by Syntology |
2026-05 (from id) |
xiaofengShi/MechVQA/evaluation/mechvqa_eval/io_utils.py 0596526ef5c6ca8e |
ran
|
Apache-2.0 (permissive) |
| SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation added by Syntology |
2026-05 (from id) |
labhai/SERC/src/utils.py 2d3cbbede5014185 |
ran
fingerprinted |
no licence file found · pointer only |
| Curriculum Learning for Safety Alignment added by Syntology |
2026-05 (from id) |
Sandeep5500/curriculum-learning-for-safety/src/phase1/create_curriculum.py 12db38ad37394041 |
ran
|
no licence file found · pointer only |
| Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning added by Syntology |
2026-05 (from id) |
AlexFanw/LegalSearch-R1/user/legalsearch_data_process.py 0eed8cf50397ec71 |
ran
|
Apache-2.0 (permissive) |
| The Model Parking Tax: Quantifying the Hidden Energy Cost of Always-On GPU Model Deployment added by Syntology |
2026-05 (from id) |
8bitai/gpu-parking-tax/analysis/generate_new_experiments_figures.py 08acd783abc2687d |
unverified |
no licence file found · pointer only |
| Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild added by Syntology |
2026-05 (from id) |
Tencent-Hunyuan/Hy-MT2/IFMTBench/run_eval.py 3b68b7b1239b923f |
ran
|
licence not identified · pointer only |
| Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training added by Syntology |
2026-05 (from id) |
cmoyacal/tie-training/src/llm/concat_for_training.py 4ea2119c8445dd18 |
ran
|
MIT (permissive) |
| LegalCiteBench: Evaluating Citation Reliability in Legal Language Models added by Syntology |
2026-05 (from id) |
Sijia711/LegalCiteBench/legal-citation-benchmark-clean/analysis/run_prompt_mitigation.py c06818718c5e5355 |
ran
|
MIT (permissive) |
| CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG added by Syntology |
2026-05 (from id) |
ideas-labo/cds4rag/Run_util.py 4561fb8f5d531d34 |
ran
|
Apache-2.0 (permissive) |
| Reproducing Complex Set-Compositional Information Retrieval added by Syntology |
2026-05 (from id) |
informagi/Complex-Set-Compositional-IR/code/cost_estimates/estimate_reranking_costs.py a5175d986e184179 |
ran
|
no licence file found · pointer only |
| Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation added by Syntology |
2026-04 (from id) |
synthiumjp/bcb-sandbagging-pilot/generate_figures.py 2467069a4724a8c1 |
ran
|
no licence file found · pointer only |
| From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling added by Syntology |
2026-04 (from id) |
CHIANGEL/Agora-Opt/code/Agora-Opt/src/debate_memory/augment_memory_from_standalone_runs.py 888f5c53e1db5aa6 |
ran
|
no licence file found · pointer only |
| From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling added by Syntology |
2026-04 (from id) |
CHIANGEL/Agora-Opt/code/Agora-Opt/src/debate_memory/debate_memory_builder.py 70016d2a0a8b58c9 |
ran
|
no licence file found · pointer only |
| SCURank: Ranking Multiple Candidate Summaries with Summary Content Units for Enhanced Summarization added by Syntology |
2026-04 (from id) |
IKMLab/SCURank/utils.py d18bc6c6c02cdbb2 |
unverified |
MIT (permissive) |
| SCURank: Ranking Multiple Candidate Summaries with Summary Content Units for Enhanced Summarization added by Syntology |
2026-04 (from id) |
IKMLab/SCURank/experiments/human-compare/data_build.py bda0f2927436562b |
unverified |
MIT (permissive) |
| LLM-Extracted Covariates for Clinical Causal Inference: Rethinking Integration Strategies added by Syntology |
2026-04 (from id) |
fpxlei/LLM-Covariates-Causal/finetune/train_lora.py 1bfdd49cdba2faac |
unverified |
MIT (permissive) |
| Benchmarking Real-Time Question Answering via Executable Code Workflows added by Syntology |
2026-04 (from id) |
leaves-slient/RT-Bench/DeepResearch/evaluation/evaluate_hle_official.py e587e0fad56e82f2 |
unverified |
no licence file found · pointer only |
| A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities added by Syntology |
2026-04 (from id) |
cjia7/DPR/src/npti/eval/eval_bbh.py 8fe69f136a96c214 |
unverified |
no licence file found · pointer only |
| A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities added by Syntology |
2026-04 (from id) |
cjia7/DPR/src/npti/eval/gpt4_score.py c7a248507261f132 |
unverified |
no licence file found · pointer only |
| Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews added by Syntology |
2026-04 (from id) |
GGLAB-KU/LOBSTER/base_runner.py 478febea332e6ec2 |
unverified |
licence not identified · pointer only |
| π 2 : Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models added by Syntology |
2026-04 (from id) |
vt-pi-squared/pi-squared/datasets/ours/code/v1_merge_and_postprocess_full_data_pipeline.py bbf4c6a9a91e3bea |
unverified |
Apache-2.0 (permissive) |
| GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces added by Syntology |
2026-04 (from id) |
ornamentt/GeoBrowse/Evaluation/baseline/evaluate.py e587e0fad56e82f2 |
unverified |
MIT (permissive) |
| Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation added by Syntology |
2026-04 (from id) |
FanZT6/OptimSyn/data_synthesis/generate_qa.py fc7b419016b1e837 |
unverified |
no licence file found · pointer only |
| Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation added by Syntology |
2026-04 (from id) |
FanZT6/OptimSyn/data_synthesis/generate_qa_eval.py 761da305e0f257c4 |
unverified |
no licence file found · pointer only |
| Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation added by Syntology |
2026-04 (from id) |
FanZT6/OptimSyn/data_synthesis/generate_rubrics.py 131f8568b132a083 |
unverified |
no licence file found · pointer only |
| HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific Literature added by Syntology |
2026-03 (from id) |
basiralab/HGNet/datasets/SPHERE/prepare_sphere.py d796bc0f8d962649 |
unverified |
MIT (permissive) |
| A Tutorial Review of Bayesian Optimization with Gaussian Processes to Accelerate Stationary Point Searches added by Syntology |
2026-03 (from id) |
lode-org/ChemGP/benchmarks/analysis/summarize.py 6924be70a5282217 |
unverified |
MIT (permissive) |
| CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents added by Syntology |
2026-03 (from id) |
dcgm/czechtopic/evaluation/common.py 0a1da25975c47643 |
unverified |
no licence file found · pointer only |
| The Wikidata Query Logs Dataset added by Syntology |
2026-02 (from id) |
ad-freiburg/wikidata-query-logs/check_overlap.py 3ad6b4253f5b5e42 |
unverified |
Apache-2.0 (permissive) |
| Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges added by Syntology |
2026-02 (from id) |
ZDCSlab/Rubrics-as-an-Attack-Surface/downstream_eval/eval/utils_pairwise.py 400eeff90e1ba381 |
unverified |
MIT (permissive) |
| Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model added by Syntology |
2026-02 (from id) |
AndrewWTY/SecCoderX/vul_induce_prompt_pipeline/inference_instructions_with_vllm.py 0348ea9288baf003 |
unverified |
no licence file found · pointer only |
| Chunky Post-Training: Data Driven Failures of Generalization added by Syntology |
2026-02 (from id) |
seoirsem/SURF/surf/core/streaming.py 0b4599e397fec376 |
unverified |
no licence file found · pointer only |
| UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos added by Syntology |
2026-01 (from id) |
aifinlab/UniFinEval/evaluate_py/data_loader.py fca995540b6041ae |
unverified |
Apache-2.0 (permissive) |
| On the Paradoxical Interference between Instruction-Following and Task Solving added by Syntology |
2026-01 (from id) |
kijlk/IF-Interference/src/math_and_qa/math_utils.py 505e67e438e9eb19 |
unverified |
no licence file found · pointer only |
| Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding added by Syntology |
2026-01 (from id) |
scheshmi/CrossStyle-MMR/train/sft_combined.py c9bf1dffc4d2016c |
unverified |
no licence file found · pointer only |
| Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding added by Syntology |
2026-01 (from id) |
ooochen-30/Emotion-LLaMA-v2/evaluation/score_split.py a118c735eadfbee8 |
unverified |
BSD-3-Clause (permissive) |
| Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding added by Syntology |
2026-01 (from id) |
ooochen-30/Emotion-LLaMA-v2/evaluation/score_split_sentiment_no_neutral.py af8fdb6ca114a86c |
unverified |
BSD-3-Clause (permissive) |
| Fast and Accurate Causal Parallel Decoding using Jacobi Forcing FAST AND ACCURATE CAUSAL PARALLEL DECODING USING JACOBI FORCING added by Syntology |
2025-12 (from id) |
hao-ai-lab/JacobiForcing/JacobiForcing/ar_inference_baseline.py 163eb1381b9a92fb |
unverified |
Apache-2.0 (permissive) |
| Base Models Know How to Reason, Thinking Models Learn When added by Syntology |
8 Oct 2025 |
cvenhoff/thinking-llms-interp/human_eval/sample.py 80fe4b7f121fc9b9 |
unverified |
no licence file found · pointer only |
| PruneCD: Contrasting Pruned Self Model to Improve Decoding Factuality added by Syntology |
2025-09 (from id) |
hoeng4/PruneCD/2_benchmark/5_strqa.py ff3b5ca99f326851 |
unverified |
no licence file found · pointer only |
| PruneCD: Contrasting Pruned Self Model to Improve Decoding Factuality added by Syntology |
2025-09 (from id) |
hoeng4/PruneCD/2_benchmark/6_gsm8k.py b5d313d6f48bb56a |
unverified |
no licence file found · pointer only |
| CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion added by Syntology |
2025-09 (from id) |
KDEGroup/CodeRAG/coderag/benchmark/common.py 08e488a1ef10ab01 |
unverified |
MIT (permissive) |
| OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft added by Syntology |
2025-09 (from id) |
CraftJarvis/OpenHA/openagents/utils/file_op.py 3ac774349fa56ea5 |
unverified |
MIT (permissive) |
| RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events added by Syntology |
2025-09 (from id) |
Bili-Sakura/RSCC/evaluation/metrics.py e7b871016c230aac |
ran · our draft was wrong
|
no licence file found · pointer only |
| WebSailor: Navigating Super-human Reasoning for Web Agent |
3 Jul 2025 |
alibaba-nlp/webwalker/evaluation/evaluate_hle_official.py e587e0fad56e82f2 |
unverified |
Apache-2.0 (permissive) |
| Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation |
3 Jun 2025 |
CuSO4-Chen/PLI/src/evaluation/gsm8k_eval.py e9a7befa1d14a877 |
ran
|
MIT (permissive) |
| SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL |
31 May 2025 |
quge2023/SHARE/src/utils.py 32da31b37eea7836 |
unverified |
Apache-2.0 (permissive) |
| Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation |
30 May 2025 |
yczhou001/longbench-t2i/utils/utils.py fdea7da6487b6a46 |
unverified |
MIT (permissive) |
| VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification |
16 May 2025 |
wilyub/verithoughts/verithoughts/multi-turn-yosys-generation.py 4398a8aef9b9aba0 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents |
18 May 2025 |
navidmdn/ESC-Judge/multidim_judge_from_merged.py 720f3f0012bdfd66 |
unverified |
Apache-2.0 (permissive) |
| ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents |
18 May 2025 |
navidmdn/ESC-Judge/multidim_geval.py 885cdf0c8392f181 |
unverified |
Apache-2.0 (permissive) |
| MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning |
15 May 2025 |
identical code first harvested elsewhere 17081a7b73a41850 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Rewriting Pre-Training Data Boosts LLM Performance in Math and Code |
5 May 2025 |
rioyokotalab/swallow-code-math/src/math/finemath-4+-rewrite-v1.py 0d9c05c05e8d3560 |
ran · our draft was wrong
|
no licence file found · pointer only |
| RWKV-X: A Linear Complexity Hybrid Language Model |
30 Apr 2025 |
howard-hou/rwkv-x/evaluation/eval_long_loss.py 933d4bd3c4bd2d14 |
unverified |
MIT (permissive) |
| Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models |
10 Apr 2025 |
lum1104/eibench/EIBench/human_eval/web_ann_basic.py e85a5515b1b5db8b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2504.04635 |
2025-04 (from id) |
patqdasilva/steering-off-course/DoLa/strqa_eval.py ff3b5ca99f326851 |
unverified |
MIT (permissive) |
| arXiv:2504.04635 |
2025-04 (from id) |
patqdasilva/steering-off-course/DoLa/gsm8k_eval.py b5d313d6f48bb56a |
unverified |
MIT (permissive) |
| TEMPLE:Temporal Preference Learning of Video LLMs via Difficulty Scheduling and Pre-SFT Alignment |
21 Mar 2025 |
lscpku/temple/preprocess.py db41886fef3918ad |
ran · our draft was wrong
|
no licence file found · pointer only |
| MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion |
20 Mar 2025 |
qizhipei/mathfusion/evaluation/dart_math/utils.py fb57a712227893bd |
ran
|
Apache-2.0 (permissive) |
| EnvBench: A Benchmark for Automated Environment Setup |
18 Mar 2025 |
JetBrains-Research/EnvBench/env_setup_utils/analysis/analysis_utils.py 7035f0d171657373 |
unverified |
MIT (permissive) |
| EnvBench: A Benchmark for Automated Environment Setup |
18 Mar 2025 |
JetBrains-Research/EnvBench/env_setup_utils/analysis/analyze_results.py 97ada5f0eb559039 |
unverified |
MIT (permissive) |
| EnvBench: A Benchmark for Automated Environment Setup |
18 Mar 2025 |
JetBrains-Research/EnvBench/env_setup_utils/analysis/scripts_viewer.py db2166cef66a916e |
unverified |
MIT (permissive) |
| VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search |
13 Mar 2025 |
tiger-ai-lab/visualwebinstruct/VisualWebInstruct/answer_alignment.py cd1b2e17d5682457 |
unverified |
MIT (permissive) |
| XIFBench: Evaluating Large Language Models on Multilingual Instruction Following |
10 Mar 2025 |
zhenyuli801/XIFBench/A1_sample_instructions.py d4c04971c3f6dc69 |
unverified |
MIT (permissive) |
| MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning |
10 Mar 2025 |
gersteinlab/medagents-benchmark/output/utils.py b07109264c29d074 |
unverified |
MIT (permissive) |
| WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation |
10 Mar 2025 |
PKU-YuanGroup/WISE/vllm_eval.py 0bc2febf8ef317cf |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation |
9 Mar 2025 |
Hritikbansal/videophy/VIDEOPHY2/data_utils/xgpt3_dataset.py f3a23711c5501382 |
ran
|
MIT (permissive) |
| HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation |
3 Mar 2025 |
0russwest0/HoH/qa_generate/src/utils.py 372d16df79db74a4 |
unverified |
Apache-2.0 (permissive) |
| Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs |
24 Feb 2025 |
emergent-misalignment/emergent-misalignment/open_models/utils.py 1d683a1b53d1615a |
unverified |
MIT (permissive) |
| Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning |
20 Feb 2025 |
Unakar/Logic-RL/eval_kk/main_eval_instruct.py 0998b08e39672d66 |
unverified |
Apache-2.0 (permissive) |
| TokenSkip: Controllable Chain-of-Thought Compression in LLMs |
17 Feb 2025 |
hemingkx/TokenSkip/LLMLingua.py 3dc03e3fd9c53b25 |
unverified |
Apache-2.0 (permissive) |
| C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation |
10 Feb 2025 |
Chen-GX/C-3PO/C-3PO/utils.py b27698eecee6ba64 |
unverified |
Apache-2.0 (permissive) |
| CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing |
4 Feb 2025 |
aiming-lab/CITER/src/pipeline/token_route.py 0bdfccf9d2a4cbc6 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Kimi k1.5: Scaling Reinforcement Learning with LLMs |
22 Jan 2025 |
mathllm/math-v/models/utils.py 17081a7b73a41850 |
ran · our draft was wrong
|
MIT (permissive) |
| Kimi k1.5: Scaling Reinforcement Learning with LLMs |
22 Jan 2025 |
mathllm/math-v/models/GPT_with_caption.py 8bcc3df886c26873 |
ran
|
MIT (permissive) |
| Kimi k1.5: Scaling Reinforcement Learning with LLMs |
22 Jan 2025 |
mathllm/math-v/models/GPT4.py f57347eff09ff0f5 |
unverified |
MIT (permissive) |
| Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems |
12 Dec 2024 |
RUCAIBox/Slow_Thinking_with_LLMs/STILL-3-TOOL/data_synthesis/coding_by_thinking.py 4b0aef908978da14 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models |
2024-12 (from id) |
shansongliu/M2UGen/DataSet/MUImage/llava_caption.py 9eb0b14fc790de41 |
ran · our draft was wrong
|
MIT (permissive) |
| MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization |
9 Dec 2024 |
aiming-lab/mmedpo/eval/eval_report.py 77678f2758141df9 |
unverified |
Apache-2.0 (permissive) |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases |
3 Dec 2024 |
kki2eve/agri-llava/agri_llava/eval/run_eval.py 77678f2758141df9 |
unverified |
Apache-2.0 (permissive) |
| SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models |
1 Nov 2024 |
JayZhang42/SLED/utils/utils_gsm8k.py 999e06e03a65eb6d |
unverified |
no licence file found · pointer only |
| SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models |
1 Nov 2024 |
JayZhang42/SLED/utils/utils_strqa.py 21ebee30ce3d7085 |
unverified |
no licence file found · pointer only |
| Can Language Models Learn to Skip Steps? |
4 Nov 2024 |
tengxiaoliu/LM_skip/src/evaluate_aoa.py 81f1e40889f24580 |
unverified |
no licence file found · pointer only |
| WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines |
16 Oct 2024 |
worldcuisines/worldcuisines/evaluation/score/score.py bf44cfc1771f959a |
ran
|
Apache-2.0 (permissive) |
| MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark |
15 Oct 2024 |
xfey/mctbench/eval/choice_stat.py 9eb0b14fc790de41 |
ran · our draft was wrong
|
licence not identified · pointer only |
| Language Imbalance Driven Rewarding for Multilingual Self-improving |
11 Oct 2024 |
ZNLP/Language-Imbalance-Driven-Rewarding/utils/utils.py 4ccfe92e7eff858a |
ran
|
no licence file found · pointer only |
| MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code |
10 Oct 2024 |
mathllm/mathcoder2/data_processing/mathematical_code/process.py 17081a7b73a41850 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code |
10 Oct 2024 |
mathllm/mathcoder2/data_processing/decontamination/exact_match_and_13-gram.py 0d0b215c235ca58d |
ran
|
no licence file found · pointer only |
| MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code |
10 Oct 2024 |
mathllm/mathcoder2/data_processing/mathematical_code/convert_to_text.py 03083657f2dac243 |
ran
|
no licence file found · pointer only |
| TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention |
7 Oct 2024 |
DerrickYLJ/TidalDecode/src/utils.py a8b4a0de765a348c |
ran
|
Apache-2.0 (permissive) |