| Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark added by Syntology |
2026-05 (from id) |
SUFE-AILAB/WSADBench/WSADBench/baseline/DualMGAN/model.py d82550e8f4e38ef6 |
unverified |
MIT (permissive) |
| MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility added by Syntology |
2026-05 (from id) |
gsasikiran/MLReplicate-benchmarking/AgentLaboratory/mlesolver.py c7d4030263209148 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning added by Syntology |
2026-02 (from id) |
zhh6425/CalibRL/src/core_algos.py ffeb684124c8f7e9 |
ran
|
no licence file found · pointer only |
| CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare added by Syntology |
2025-12 (from id) |
AikyamLab/clinic/evaluation/colloquial/colloquial_eval.py 7fdf75bfa127d457 |
unverified |
MIT (permissive) |
| CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare added by Syntology |
2025-12 (from id) |
AikyamLab/clinic/evaluation/disparagement/disp_eval.py a5fb7b464744eb20 |
unverified |
MIT (permissive) |
| CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare added by Syntology |
2025-12 (from id) |
AikyamLab/clinic/evaluation/exaggerated_safety/exagsaf.py 987ba214da993053 |
unverified |
MIT (permissive) |
| ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind |
29 May 2025 |
ulab-uiuc/ToMAP/verl/env_feedback/argument_graph.py a8c0b5b7b7144100 |
unverified |
Apache-2.0 (permissive) |
| AgentRxiv: Towards Collaborative Autonomous Research |
23 Mar 2025 |
samuelschmidgall/agentlaboratory/mlesolver.py c7d4030263209148 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| AgentRxiv: Towards Collaborative Autonomous Research |
23 Mar 2025 |
samuelschmidgall/agentlaboratory/agents.py 3c76d8364248f49a |
unverified |
MIT (permissive) |
| Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity |
2 Mar 2025 |
hypasd-art/ETAPP/evaluation/evaluate.py 7ab562f4ba9b1d82 |
unverified |
no licence file found · pointer only |
| DriveMLLM: A Benchmark for Spatial Understanding with Multimodal Large Language Models in Autonomous Driving |
20 Nov 2024 |
xiandaguo/drive-mllm/evaluation/eval_from_json.py bbd19315c49635ff |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/audio_duration_prediction.py 67c5af6a71776bab |
unverified |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/audio_editing_identification.py ab963e47043f8825 |
unverified |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/audio_spatial_distance.py 9e34dbd0951e6bdc |
unverified |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/code_switching_count.py 35daeec0cf300934 |
unverified |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/exact_match.py ef715545375e59c0 |
unverified |
no licence file found · pointer only |
| Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks |
8 Nov 2024 |
dynamic-superb/dynamic-superb/api/metrics/llm_classification.py 31d6756ad131ace0 |
unverified |
no licence file found · pointer only |
| DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution |
4 Nov 2024 |
yueyang130/DeeR-VLA/bayesian_optimization.py 1fff25335fef27a2 |
unverified |
Apache-2.0 (permissive) |
| What If the Input is Expanded in OOD Detection? |
24 Oct 2024 |
tmlr-group/CoVer/DNNs/ash.py f03fa80d09174dd7 |
unverified |
no licence file found · pointer only |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/inference/utils.py 5321a2aa0d7681cb |
unverified |
no licence file found · pointer only |
| What makes your model a low-empathy or warmth person: Exploring the Origins of Personality in LLMs |
7 Oct 2024 |
kaustpradalab/LLM-Persona-Steering/src/steer_experiments/RepE/analysis.py 6d60c2f13b005719 |
ran
|
no licence file found · pointer only |
| What makes your model a low-empathy or warmth person: Exploring the Origins of Personality in LLMs |
7 Oct 2024 |
kaustpradalab/LLM-Persona-Steering/src/steer_experiments/SAE/analysis_origin.py 56ea01da64a80821 |
ran
|
no licence file found · pointer only |
| Transferring disentangled representations: bridging the gap between synthetic and real images |
26 Sep 2024 |
JacopoDapueto/transfer_disentanglement/src/evaluation/metrics/omes.py bbe86af17ad979c9 |
unverified |
Apache-2.0 (permissive) |
| Zero-Shot Detection of LLM-Generated Text using Token Cohesiveness |
25 Sep 2024 |
Shixuan-Ma/TOCSIN/TOCSIN.py 9c87e2fd2309802e |
unverified |
MIT (permissive) |
| Archon: An Architecture Search Framework for Inference-Time Techniques |
23 Sep 2024 |
scalingintelligence/archon/src/archon/benchmarks/arena_hard_auto/gen_judgment.py b2d83f0b138fe334 |
unverified |
Apache-2.0 (permissive) |
| Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models |
18 Sep 2024 |
value4ai/gpv/gpv/utils.py 07996379f9e365b6 |
ran
fingerprinted |
MIT (permissive) |
| CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models |
4 Sep 2024 |
ecnu-icalk/educhat-math/evaluation/gpt-4o-score_evaluation.py ac28ae0c3a3f6b14 |
ran
|
no licence file found · pointer only |
| ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO |
17 Jun 2024 |
snumprlab/SRT/inference/utils.py 5321a2aa0d7681cb |
unverified |
no licence file found · pointer only |
| VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text |
10 Jun 2024 |
tianyu-z/vcr/src/evaluation/gather_results.py 9b842bc0490f11b4 |
ran
|
CC-BY-SA-4.0 (copyleft) · pointer only |
| Eliciting Informative Text Evaluations with Large Language Models |
23 May 2024 |
yx-lu/Eliciting-Informative-Text-Evaluations-with-Large-Language-Models/experiments/baseline.py fc5eb58baae10b8a |
unverified |
CC-BY-4.0 · pointer only |
| Is Your LLM Outdated? Evaluating LLMs at Temporal Generalization |
14 May 2024 |
freedomintelligence/freshbench/script/score.py b21ef2f2569d9b67 |
ran
fingerprinted |
no licence file found · pointer only |
| Continuous Language Model Interpolation for Dynamic and Controllable Text Generation |
10 Apr 2024 |
skangasl/continuous-lm-interpolation/evaluation/create_plots.py 72a3e012682bb367 |
ran
|
no licence file found · pointer only |
| Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning |
25 Mar 2024 |
deepcs233/visual-cot/llava/eval/eval_cot_score.py 5b700f36a74a7df5 |
ran
|
Apache-2.0 (permissive) |
| RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations |
27 Feb 2024 |
explanare/ravel/src/methods/linear_adversarial_probe.py 4637704289ff6a8a |
ran
|
MIT (permissive) |
| When LLMs Meet Cunning Texts: A Fallacy Understanding Benchmark for Large Language Models |
16 Feb 2024 |
thukelab/flub/code/analysis.py 28af50493dc61691 |
ran · honoured contract
|
licence not identified · pointer only |
| Escalation Risks from Language Models in Military and Diplomatic Decision-Making |
7 Jan 2024 |
jprivera44/EscalAItion/manual_evaluation.py 74954cd7a1260147 |
ran
|
no licence file found · pointer only |
| HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models |
21 Dec 2023 |
picsart-ai-research/hd-painter/metrics/aesthetic.py e670aa8de780268c |
unverified |
MIT (permissive) |
| Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer |
28 Nov 2023 |
diffusion-motion-transfer/diffusion-motion-transfer/motion_fidelity_score.py 16286335dc3aec83 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Metric Space Magnitude for Evaluating the Diversity of Latent Representations |
27 Nov 2023 |
renata-turkes/turkevs2022on/SRC/model.py 30e11bb16f36a080 |
ran
|
no licence file found · pointer only |
| Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models |
19 Nov 2023 |
Docta-ai/docta/docta/core/knn.py 7e1d6ecb79e8a1bc |
ran
|
licence not identified · pointer only |
| Instruct and Extract: Instruction Tuning for On-Demand Information Extraction |
24 Oct 2023 |
yzjiao/on-demand-ie/evaluation/rougel_for_content.py 052514aef99856ca |
ran
|
no licence file found · pointer only |
| Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization |
23 Oct 2023 |
llm-eff/FedPepTAO/decoder-only-gpt2/get_score.py 7b1aa9397d2f3a55 |
unverified |
no licence file found · pointer only |
| Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization |
23 Oct 2023 |
llm-eff/FedPepTAO/decoder-only-llama/get_score.py 48cf69dafa7e622e |
unverified |
no licence file found · pointer only |
| Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization |
23 Oct 2023 |
llm-eff/FedPepTAO/encoder-only-roberta-large/get_score.py 822768feef873f11 |
unverified |
no licence file found · pointer only |
| HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models |
23 Oct 2023 |
zli12321/qa_metrics/qa_metrics/RewardBert.py bb06d405aa0fd21d |
unverified |
MIT (permissive) |
| ExpertQA: Expert-Curated Questions and Attributed Answers |
14 Sep 2023 |
chaitanyamalaviya/expertqa/modeling/fact_score/factscore.py 6c4b02266edc91ba |
ran · violated contract
|
MIT (permissive) |
| CharacterChat: Learning towards Conversational AI with Personalized Social Support |
20 Aug 2023 |
morecry/characterchat/model/demo/chat_demo.py 5efb4305d6662763 |
ran
|
no licence file found · pointer only |
| Uni-NLX: Unifying Textual Explanations for Vision and Vision-Language Tasks |
17 Aug 2023 |
fawazsammani/nlxgpt/explain_predict/ep_vqaX.py 04a66156efbb1777 |
ran
fingerprinted |
no licence file found · pointer only |
| FunQA: Towards Surprising Video Comprehension |
26 Jun 2023 |
jingkang50/funqa/gpt4_eval.py b1ee5d8a47a7d92a |
unverified |
MIT (permissive) |
| LNL+K: Enhancing Learning with Noisy Labels Through Noise Source Knowledge Integration |
20 Jun 2023 |
sunnysiqi/lnl_k/adaptation_methods/fine_k.py 56ccb4fcec6f5a98 |
unverified |
MIT (permissive) |
| Large Language Models of Code Fail at Completing Code with Potential Bugs |
6 Jun 2023 |
amazon-science/buggy-code-completion/src/infiller/infill_line.py 5f28bf9a2c267875 |
unverified |
Apache-2.0 (permissive) |
| Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage |
22 May 2023 |
hanyins/lm_association_quantification/LAMA/analysis.py 54f850b6751cfb2e |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Unsupervised Entity Alignment for Temporal Knowledge Graphs |
1 Feb 2023 |
zju-daily/dualmatch/processing.py 68824ba6e717496e |
unverified |
Apache-2.0 (permissive) |
| Functional Output Regression with Infimal Convolution: Exploring the Huber and $ε$-insensitive Losses |
16 Jun 2022 |
allambert/foreg/model_selection/model_selection.py 00e351b580b1143b |
unverified |
MIT (permissive) |
| Diverse Weight Averaging for Out-of-Distribution Generalization |
19 May 2022 |
alexrame/diwa/domainbed/lib/misc.py 837d97171609bbbf |
unverified |
Apache-2.0 (permissive) |
| Out-of-Distribution Detection with Deep Nearest Neighbors |
13 Apr 2022 |
deeplearning-wisc/knn-ood/util/score.py b874b3a0ce06fd5f |
unverified |
no licence file found · pointer only |
| Linear Adversarial Concept Erasure |
28 Jan 2022 |
shauli-ravfogel/rlace-icml/rlace.py d7edaf170e8ed046 |
unverified |
no licence file found · pointer only |
| Linear Adversarial Concept Erasure |
28 Jan 2022 |
shauli-ravfogel/rlace-icml/rlace.py b765918151ddbd4e |
unverified |
no licence file found · pointer only |
| Linear Adversarial Concept Erasure |
28 Jan 2022 |
shauli-ravfogel/adv-kernel-removal/relaxed_inlp.py 2e9f7545f8882b0b |
unverified |
no licence file found · pointer only |
| FastFlow: Unsupervised Anomaly Detection and Localization via 2D Normalizing Flows |
15 Nov 2021 |
RistoranteRist/FastFlow/utils.py 2f2d58311dd3969d |
unverified |
MIT (permissive) |
| Ensemble and Auxiliary Tasks for Data-Efficient Deep Reinforcement Learning |
5 Jul 2021 |
NUS-LID/RENAULT/benchmark.py 640097e0a5b3d620 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| FINE Samples for Learning with Noisy Labels |
23 Feb 2021 |
Kthyeon/FINE_official/dynamic_selection/selection/svd_classifier.py 1b10ac39415f471f |
ran · fixture could not drive it
|
no licence file found · pointer only |
| RLCard: A Toolkit for Reinforcement Learning in Card Games |
10 Oct 2019 |
datamllab/rlcard/rlcard/envs/blackjack.py 6c08fd179035271c |
ran · honoured contract
|
MIT (permissive) |
| Image Synthesis From Reconfigurable Layout and Style |
20 Aug 2019 |
stanifrolov/attrlostgan/eval/attr_f1.py f19ef3041e7db2b2 |
ran · honoured contract
|
no licence file found · pointer only |
| End-to-End Task-Completion Neural Dialogue Systems |
3 Mar 2017 |
AtmaHou/UserSimulator/src/collect_result.py cdefc0c373bb5645 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields |
24 Nov 2016 |
ibm/max-human-pose-estimator/core/tf_pose/pafprocess/pafprocess.py 418d24f644276625 |
unverified |
Apache-2.0 (permissive) |
| arXiv:aaai_34690 |
|
ZongyueQin/DSBD/evaluation.py 214e835658c1d3c9 |
unverified |
MIT (permissive) |
| arXiv:2025.findings-acl.931 |
|
ikergarcia1996/T-Projection/label_projection.py fafc9230eec0ad80 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.findings-acl.1384 |
|
Muennighoff/sgpt/crossencoder/beir/openai_search_endpoint_functionality.py e71e0d710acd1e64 |
unverified |
MIT (permissive) |