| The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction added by Syntology |
2026-09 (from id) |
hopit-ai/Moda_ner/suite/text/score.py 7d080e2859a10e18 |
unverified |
MIT (permissive) |
| REVES: REvision and VErification-Augmented Training for Test-Time Scaling added by Syntology |
2026-06 (from id) |
yxliu02/REVES/evaluator/eval_travelplanner.py 2dddcb48090ef99f |
ran
|
no licence file found · pointer only |
| PAC-Chernoff Bounds: Understanding Generalization in the Interpolation Regime added by Syntology |
2026-06 (from id) |
Ludvins/FixedMeanGaussianProcesses/bayesipy/utils/metrics.py 9110c7d87661f572 |
ran
|
no licence file found · pointer only |
| ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation added by Syntology |
2026-05 (from id) |
zhongkaiyu/ChipMATE/paper_repro/chipbench/chipbench_score.py dd0f65afa280200f |
ran
|
licence not identified · pointer only |
| ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation added by Syntology |
2026-05 (from id) |
zhongkaiyu/ChipMATE/paper_repro/chipbench/chipbench_framework.py c6869481cba9dde6 |
unverified |
licence not identified · pointer only |
| ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation added by Syntology |
2026-05 (from id) |
zhongkaiyu/ChipMATE/paper_repro/rtllm/rtllm_framework.py a222280252a2db47 |
unverified |
licence not identified · pointer only |
| Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum Akuen Akoi Deng ⋆[0009-0007-6228-3340] , Eimantas Butkus ⋆[0009-0001-5647-0779] , Alfreds Lapkovskis [0009-0003-4424-949X] , and added by Syntology |
2026-05 (from id) |
Akuien/DNN-partitioning-and-Oflloading-framework-REAP/adaptive-framework/split_infer.py 72847bdb72a3845e |
ran · honoured contract
|
MIT (permissive) |
| A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis added by Syntology |
2026-05 (from id) |
codelucas/newspaper/newspaper/nlp.py 491f58b8741fe8d7 |
unverified |
MIT (permissive) |
| Medical Reasoning with Large Language Models: A Survey and MR-Bench added by Syntology |
2026-04 (from id) |
RXH04-USTC/Medical-Reasoning-Survey/src_eval/scorer.py abe344037ab0d857 |
ran · our draft was wrong
|
licence not identified · pointer only |
| EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning added by Syntology |
2026-01 (from id) |
EverMind-AI/EverMemOS/benchmarks/metrics/core.py 0fcb7854e27f1cad |
unverified |
Apache-2.0 (permissive) |
| EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning added by Syntology |
2026-01 (from id) |
EverMind-AI/EverMemOS/benchmarks/metrics/ir.py f05168abaac86222 |
unverified |
Apache-2.0 (permissive) |
| TROVE: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models added by Syntology |
2025-12 (from id) |
Stanford-AIMI/TRoVe/src/trove/score.py 10b39a57b20324c8 |
unverified |
MIT (permissive) |
| REALM: Recursive Relevance Modeling for LLM-based Document Re-Ranking * added by Syntology |
2025-08 (from id) |
Joeyw02/REALM/src/algorithm.py 051a1bed8cba2a19 |
unverified |
no licence file found · pointer only |
| arXiv:2507.00322 |
2025-07 (from id) |
EleutherAI/pythia/predictable-memorization/eval_memorization.py 81bc28dc8aceb9c2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning |
23 Jun 2025 |
thudm/longwriter/evaluation/eval_length.py 653c358686bbb406 |
unverified |
Apache-2.0 (permissive) |
| CausalDynamics: A large-scale benchmark for structural discovery of dynamical causal models |
22 May 2025 |
kausable/CausalDynamics/src/causaldynamics/score.py fe027f1e8d858d64 |
unverified |
MIT (permissive) |
| Deep Koopman operator framework for causal discovery in nonlinear dynamical systems |
20 May 2025 |
juannat7/kausal/kausal/utils.py f1286fc26e8a683a |
unverified |
MIT (permissive) |
| Evaluating the Goal-Directedness of Large Language Models |
16 Apr 2025 |
Crista23/goal_directedness_llms/tasks/full_task.py bb22736701cd8662 |
unverified |
Apache-2.0 (permissive) |
| Correlation and Navigation in the Vocabulary Key Representation Space of Language Models |
3 Oct 2024 |
KomeijiForce/KeyNavi/zerogen.py 19516e6551d27215 |
unverified |
no licence file found · pointer only |
| Evidence Is All You Need: Ordering Imaging Studies via Language Model Alignment with the ACR Appropriateness Criteria |
27 Sep 2024 |
michael-s-yao/radGPT/radgpt/utils.py f69f9d0af8972175 |
ran
|
MIT (permissive) |
| Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild |
17 Jul 2024 |
nicolas-richet/feature-vs-text-compound-emotion/model.py fc5685919b105cb5 |
ran
|
no licence file found · pointer only |
| Shape Arithmetic Expressions: Advancing Scientific Discovery Beyond Closed-Form Equations |
15 Apr 2024 |
krzysztof-kacprzyk/shares/experiments/benchmarks.py cb821c0bc00a2598 |
ran
|
no licence file found · pointer only |
| Satellite Imagery and AI: A New Era in Ocean Conservation, from Research to Deployment and Impact |
6 Dec 2023 |
allenai/vessel-detection-sentinels/src/training/metric.py eb4e28dc5f3a8176 |
ran
|
Apache-2.0 (permissive) |
| Fast and Accurate Factual Inconsistency Detection Over Long Documents |
19 Oct 2023 |
asappresearch/scale-score/scale_score/scorer.py 0d59d2ed3cc16146 |
unverified |
no licence file found · pointer only |
| Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model |
19 Oct 2023 |
iabhijith/bias-causal-analysis/evaluation/blimp.py e7c03f0573dfb130 |
unverified |
no licence file found · pointer only |
| Fine-Tuning LLaMA for Multi-Stage Text Retrieval |
12 Oct 2023 |
texttron/tevatron/src/tevatron/eval/metrics.py 5eb05d1d51ec1fa5 |
ran
|
Apache-2.0 (permissive) |
| Entropy-MCMC: Sampling from Flat Basins with Ease |
9 Oct 2023 |
lblaoke/emcmc/metric.py 76887d631d72b05f |
unverified |
no licence file found · pointer only |
| StoryBench: A Multifaceted Benchmark for Continuous Story Visualization |
22 Aug 2023 |
google/storybench/metrics/vtm_clip.py b482b6a0ac227ae9 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| EduChat: A Large-Scale Language Model-based Chatbot System for Intelligent Education |
5 Aug 2023 |
icalk-nlp/educhat/demo/score_utils.py 097528b8a72d71a7 |
ran
|
no licence file found · pointer only |
| Improving Retrieval-Augmented Large Language Models via Data Importance Learning |
6 Jul 2023 |
amsterdata/ragbooster/python/ragbooster/core.py cfd204a25911c17c |
unverified |
Apache-2.0 (permissive) |
| Emergent and Predictable Memorization in Large Language Models |
21 Apr 2023 |
eleutherai/pythia/predictable-memorization/eval_memorization.py 81bc28dc8aceb9c2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Trustworthy Social Bias Measurement |
20 Dec 2022 |
rishibommasani/biasmeasures/convergent_validity.py 0569862de44357f1 |
unverified |
MIT (permissive) |
| Generative Language Models for Paragraph-Level Question Generation |
8 Oct 2022 |
asahi417/lm-question-generation/lmqg/automatic_evaluation_tool/bert_score/score.py a57915bf6208b453 |
ran · honoured contract
|
MIT (permissive) |
| Embarrassingly Easy Document-Level MT Metrics: How to Convert Any Pretrained Metric Into a Document-Level Metric |
27 Sep 2022 |
amazon-science/doc-mt-metrics/bert_score/bert_score/score.py efd9b25cfcd27f97 |
unverified |
Apache-2.0 (permissive) |
| Recovering Private Text in Federated Learning of Language Models |
17 May 2022 |
princeton-sysml/film/reorder.py b902d2253e571934 |
unverified |
CC0-1.0 (permissive) |
| On the Evaluation Metrics for Paraphrase Generation |
17 Feb 2022 |
shadowkiller33/parascore/bert_score/score.py 8f1719ca230ca245 |
unverified |
Apache-2.0 (permissive) |
| WaveFake: A Data Set to Facilitate Audio Deepfake Detection |
4 Nov 2021 |
rub-syssec/wavefake/dfadetect/models/gaussian_mixture_model.py 371c6c7eedd45605 |
unverified |
MIT (permissive) |
| Explaining deep learning of galaxy morphology with saliency mapping |
2021-10 (from id) |
prabhbhambra13/xai_bar_lengths/generate_heatmaps.py 296ac592eba01d02 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Finding a Balanced Degree of Automation for Summary Evaluation |
23 Sep 2021 |
ZhangShiyue/Lite2-3Pyramid/metric/score.py 56276e9a4bd27b29 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Studying word order through iterative shuffling |
10 Sep 2021 |
malkin1729/ibis/code/ibis.py dea9d346ab79356f |
unverified |
no licence file found · pointer only |
| Pre-train or Annotate? Domain Adaptation with a Constrained Budget |
10 Sep 2021 |
allenai/kb/bin/tacred_scorer.py 9cb44938e773d88b |
unverified |
Apache-2.0 (permissive) |
| Truth Discovery in Sequence Labels from Crowds |
9 Sep 2021 |
nasimisu/truth-discovery-in-sequence-labels-from-crowds/calculations.py 35705fd6f9774602 |
ran · honoured contract
|
no licence file found · pointer only |
| BERTTune: Fine-Tuning Neural Machine Translation with BERTScore |
4 Jun 2021 |
ijauregiCMCRC/fairseq-bert-loss/fairseq/bert_score/score.py 0a91572da431f7d0 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| When is Memorization of Irrelevant Training Data Necessary for High-Accuracy Learning? |
11 Dec 2020 |
gavinrbrown1/training-data-memorization/attacks.py f80d3b29688c4d7a |
ran · our draft was wrong
|
no licence file found · pointer only |
| Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music |
2020-10 (from id) |
Dsqvival/hierarchical-structure-analysis/preprocessing/key_finding.py e46c2500ec92d4a2 |
unverified |
MIT (permissive) |
| Unsupervised Evaluation of Interactive Dialog with DialoGPT |
23 Jun 2020 |
shikib/fed/fed.py e5b4d588ad5ce486 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| AOWS: Adaptive and optimal network width search with latency constraints |
21 May 2020 |
bermanmaxim/AOWS/viterbi.py dfc753382f18458e |
unverified |
MIT (permissive) |
| Exploiting Sentence Order in Document Alignment |
30 Apr 2020 |
facebookresearch/LASER/source/mine_bitexts.py ce7491b4961154f3 |
ran
|
licence not identified · pointer only |
| Attention Guided Graph Convolutional Networks for Relation Extraction |
18 Jun 2019 |
Cartus/AGGCN_TACRED/utils/scorer.py 9cb44938e773d88b |
unverified |
MIT (permissive) |
| Attention Guided Graph Convolutional Networks for Relation Extraction |
18 Jun 2019 |
Cartus/AGGCN_TACRED/semeval/utils/scorer.py c4b42e34eb5cc2e9 |
unverified |
MIT (permissive) |
| Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View |
6 Jun 2019 |
zhuohan123/macaron-net/bert/macaron-scripts/bert/concat_short_sentences.py 43404249e690d259 |
unverified |
BSD-3-Clause (permissive) |
| An Empirical Model of Large-Batch Training |
14 Dec 2018 |
davidandym/task-conflict-in-text-to-text-learners/src/deca_metrics.py e81a8138a9405edf |
unverified |
MIT (permissive) |
| The Natural Language Decathlon: Multitask Learning as Question Answering |
20 Jun 2018 |
salesforce/decaNLP/metrics.py e81a8138a9405edf |
unverified |
BSD-3-Clause (permissive) |
| Kernel Conditional Exponential Family |
15 Nov 2017 |
MichaelArbel/KCEF/KCEF/functions.py 841efea8b489cbff |
unverified |
BSD-3-Clause (permissive) |
| Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems |
7 Aug 2015 |
mrcmoresi/sc-lstm/run_woz3.py faf421c0fe195b81 |
ran · honoured contract
|
no licence file found · pointer only |
| Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems |
7 Aug 2015 |
andy194673/nlg-sclstm-multiwoz/run_woz3.py 2d9e4c619f62021f |
ran · honoured contract
|
MIT (permissive) |
| arXiv:2024.findings-acl.848 |
|
CAS-SIAT-ConsistencyAI/NUMCoT/src/code/experience_first/run_medium.py 205a0d52ead816d9 |
unverified |
CC0-1.0 (permissive) |
| arXiv:2024.findings-acl.848 |
|
CAS-SIAT-ConsistencyAI/NUMCoT/src/code/experience_second/run_unit_measurement.py a06a8861b5c1c21e |
unverified |
CC0-1.0 (permissive) |
| arXiv:2023.findings-acl.807 |
|
amueller/word_cloud/wordcloud/tokenization.py f23a61928a841fa0 |
unverified |
MIT (permissive) |
| arXiv:2021.findings-emnlp.152 |
|
ruidan/IMN-E2E-ABSA/code/evaluation.py c293c9d2e5e4fde1 |
unverified |
Apache-2.0 (permissive) |