Browse State-of-the-Art › Benchmarking › Papers, page 42
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 42 of 56: papers 4,101 to 4,200 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers28 Nov 2023 0 repositories listed
-
Comprehensive Benchmarking of Entropy and Margin Based Scoring Metrics for Data Selection27 Nov 2023 0 repositories listed
-
FakeWatch ElectionShield: A Benchmarking Framework to Detect Fake News for Credible US Elections27 Nov 2023 0 repositories listed
-
Lightly Weighted Automatic Audio Parameter Extraction for the Quality Assessment of Consensus Auditory-Perceptual Evaluation of Voice27 Nov 2023 0 repositories listed
-
Syn3DWound: A Synthetic Dataset for 3D Wound Bed Analysis27 Nov 2023 0 repositories listed
-
ASI: Accuracy-Stability Index for Evaluating Deep Learning Models26 Nov 2023 0 repositories listed
-
Benchmarking Large Language Model Volatility26 Nov 2023 0 repositories listed
-
An Empirical Investigation into Benchmarking Model Multiplicity for Trustworthy Machine Learning: A Case Study on Image Classification24 Nov 2023 0 repositories listed
-
Large Language Models as Automated Aligners for benchmarking Vision-Language Models24 Nov 2023 0 repositories listed
-
Automated 3D Tumor Segmentation using Temporal Cubic PatchGAN (TCuP-GAN)23 Nov 2023 0 repositories listed
-
Benchmarking Toxic Molecule Classification using Graph Neural Networks and Few Shot Learning22 Nov 2023 0 repositories listed
-
Benchmarking bias: Expanding clinical AI model card to incorporate bias reporting of social and non-social factors21 Nov 2023 0 repositories listed
-
Deep State-Space Model for Predicting Cryptocurrency Price21 Nov 2023 0 repositories listed
-
Demonstrating Almost Linear Time Complexity of Bus Admittance Matrix-Based Distribution Network Power Flow: An Empirical Approach20 Nov 2023 0 repositories listed
-
Holistic Inverse Rendering of Complex Facade via Aerial 3D Scanning20 Nov 2023 0 repositories listed
-
Segment Together: A Versatile Paradigm for Semi-Supervised Medical Image Segmentation20 Nov 2023 0 repositories listed
-
Benchmarking Feature Extractors for Reinforcement Learning-Based Semiconductor Defect Localization18 Nov 2023 0 repositories listed
-
Benchmarking Machine Learning Models for Quantum Error Correction18 Nov 2023 0 repositories listed
-
Predicting the Probability of Collision of a Satellite with Space Debris: A Bayesian Machine Learning Approach17 Nov 2023 0 repositories listed
-
Domain Aligned CLIP for Few-shot Classification15 Nov 2023 0 repositories listed
-
Model Agnostic Explainable Selective Regression via Uncertainty Estimation15 Nov 2023 0 repositories listed
-
Social Bias Probing: Fairness Benchmarking for Language Models15 Nov 2023 0 repositories listed
-
Benchmarking Individual Tree Mapping with Sub-meter Imagery14 Nov 2023 0 repositories listed
-
Uncertainty estimation of machine learning spatial precipitation predictions from satellite data13 Nov 2023 0 repositories listed
-
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks13 Nov 2023 0 repositories listed
-
The Disagreement Problem in Faithfulness Metrics13 Nov 2023 0 repositories listed
-
Identification of vortex in unstructured mesh with graph neural networks11 Nov 2023 0 repositories listed
-
SeaTurtleID2022: A long-span dataset for reliable sea turtle re-identification9 Nov 2023 0 repositories listed
-
An efficiency analysis of Spanish airports8 Nov 2023 0 repositories listed
-
Prompt Sketching for Large Language Models8 Nov 2023 0 repositories listed
-
Benchmarking Deep Facial Expression Recognition: An Extensive Protocol with Balanced Dataset in the Wild6 Nov 2023 0 repositories listed
-
Benchmarking Differential Evolution on a Quantum Simulator6 Nov 2023 0 repositories listed
-
Exploitation-Guided Exploration for Semantic Embodied Navigation6 Nov 2023 0 repositories listed
-
Benchmarking a Benchmark: How Reliable is MS-COCO?5 Nov 2023 0 repositories listed
-
Learning Disentangled Speech Representations4 Nov 2023 0 repositories listed
-
An Empirical Study of Benchmarking Chinese Aspect Sentiment Quad Prediction3 Nov 2023 0 repositories listed
-
Investigating Deep-Learning NLP for Automating the Extraction of Oncology Efficacy Endpoints from Scientific Literature3 Nov 2023 0 repositories listed
-
Use of Deep Neural Networks for Uncertain Stress Functions with Extensions to Impact Mechanics3 Nov 2023 0 repositories listed
-
Decentralized Federated Learning on the Edge over Wireless Mesh Networks2 Nov 2023 0 repositories listed
-
Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs1 Nov 2023 0 repositories listed
-
SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization1 Nov 2023 0 repositories listed
-
A Two-Step Framework for Multi-Material Decomposition of Dual Energy Computed Tomography from Projection Domain31 Oct 2023 0 repositories listed
-
Next-generation MRD assays: do we have the tools to evaluate them properly?31 Oct 2023 0 repositories listed
-
Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests31 Oct 2023 0 repositories listed
-
UAV Immersive Video Streaming: A Comprehensive Survey, Benchmarking, and Open Challenges31 Oct 2023 0 repositories listed
-
A Metadata-Driven Approach to Understand Graph Neural Networks30 Oct 2023 0 repositories listed
-
Domain Generalization in Computational Pathology: Survey and Guidelines30 Oct 2023 0 repositories listed
-
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection29 Oct 2023 0 repositories listed
-
On General Language Understanding27 Oct 2023 0 repositories listed
-
ConDefects: A New Dataset to Address the Data Leakage Concern for LLM-based Fault Localization and Program Repair25 Oct 2023 0 repositories listed
-
Quantum Long Short-Term Memory (QLSTM) vs Classical LSTM in Time Series Forecasting: A Comparative Study in Solar Power Forecasting25 Oct 2023 0 repositories listed
-
RDBench: ML Benchmark for Relational Databases25 Oct 2023 0 repositories listed
-
Analyzing Multilingual Competency of LLMs in Multi-Turn Instruction Following: A Case Study of Arabic23 Oct 2023 0 repositories listed
-
A Quantitative Evaluation of Dense 3D Reconstruction of Sinus Anatomy from Monocular Endoscopic Video22 Oct 2023 0 repositories listed
-
MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation21 Oct 2023 0 repositories listed Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Standardised workflow for mass spectrometry-based single-cell proteomics data processing and analysis using the scp package20 Oct 2023 0 repositories listed
-
Almost Equivariance via Lie Algebra Convolutions19 Oct 2023 0 repositories listed
-
Benchmarking GPUs on SVBRDF Extractor Model19 Oct 2023 0 repositories listed
-
Alexpaca: Learning Factual Clarification Question Generation Without Examples17 Oct 2023 0 repositories listed
-
A Novel Benchmarking Paradigm and a Scale- and Motion-Aware Model for Egocentric Pedestrian Trajectory Prediction16 Oct 2023 0 repositories listed
-
An Empirical Study of Super-resolution on Low-resolution Micro-expression Recognition16 Oct 2023 0 repositories listed
-
Assessing Encoder-Decoder Architectures for Robust Coronary Artery Segmentation16 Oct 2023 0 repositories listed
-
BanglaNLP at BLP-2023 Task 1: Benchmarking different Transformer Models for Violence Inciting Text Detection in Bengali16 Oct 2023 0 repositories listed
-
Evaluating Robustness of Visual Representations for Object Assembly Task Requiring Spatio-Geometrical Reasoning15 Oct 2023 0 repositories listed
-
Prompting Scientific Names for Zero-Shot Species Recognition15 Oct 2023 0 repositories listed
-
Benchmarking the Sim-to-Real Gap in Cloth Manipulation14 Oct 2023 0 repositories listed
-
A Benchmarking Protocol for SAR Colorization: From Regression to Deep Learning Approaches12 Oct 2023 0 repositories listed
-
Investigating the Robustness and Properties of Detection Transformers (DETR) Toward Difficult Images12 Oct 2023 0 repositories listed
-
Who Said That? Benchmarking Social Media AI Detection12 Oct 2023 0 repositories listed
-
Deep Reinforcement Learning for Autonomous Cyber Defence: A Survey11 Oct 2023 0 repositories listed
-
FedSym: Unleashing the Power of Entropy for Benchmarking the Algorithms for Federated Learning11 Oct 2023 0 repositories listed
-
Hypergraph Neural Networks through the Lens of Message Passing: A Common Perspective to Homophily and Architecture Design11 Oct 2023 0 repositories listed
-
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms11 Oct 2023 0 repositories listed
-
Risk Aware Benchmarking of Large Language Models11 Oct 2023 0 repositories listed
-
CAFA-evaluator: A Python Tool for Benchmarking Ontological Classification Methods10 Oct 2023 0 repositories listed
-
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets10 Oct 2023 0 repositories listed
-
Distributed Evolution Strategies with Multi-Level Learning for Large-Scale Black-Box Optimization9 Oct 2023 0 repositories listed
-
Benchmarking Large Language Models with Augmented Instructions for Fine-grained Information Extraction8 Oct 2023 0 repositories listed
-
Beyond Text: A Deep Dive into Large Language Models' Ability on Understanding Graph Data7 Oct 2023 0 repositories listed
-
Bringing Quantum Algorithms to Automated Machine Learning: A Systematic Review of AutoML Frameworks Regarding Extensibility for QML Algorithms6 Oct 2023 0 repositories listed
-
6 Oct 2023 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Full-scale modal testing of a Hawk T1A aircraft for benchmarking vibration-based methods6 Oct 2023 0 repositories listed
-
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation6 Oct 2023 0 repositories listed
-
Profit: Benchmarking Personalization and Robustness Trade-off in Federated Prompt Tuning6 Oct 2023 0 repositories listed
-
A Review of Deep Reinforcement Learning in Serverless Computing: Function Scheduling and Resource Auto-Scaling5 Oct 2023 0 repositories listed
-
Benchmarking a foundation LLM on its ability to re-label structure names in accordance with the AAPM TG-263 report5 Oct 2023 0 repositories listed
-
Deep Reinforcement Learning Algorithms for Hybrid V2X Communication: A Benchmarking Study4 Oct 2023 0 repositories listed
-
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference4 Oct 2023 0 repositories listed
-
On the Performance of Multimodal Language Models4 Oct 2023 0 repositories listed
-
Benchmarking and Improving Generator-Validator Consistency of Language Models3 Oct 2023 0 repositories listed
-
EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods3 Oct 2023 0 repositories listed
-
3 Oct 2023 0 repositories listed
-
A New Real-World Video Dataset for the Comparison of Defogging Algorithms2 Oct 2023 0 repositories listed
-
CoDBench: A Critical Evaluation of Data-driven Models for Continuous Dynamical Systems2 Oct 2023 0 repositories listed
-
Adaptive Control of an Inverted Pendulum by a Reinforcement Learning-based LQR Method30 Sep 2023 0 repositories listed
-
The Sparsity Roofline: Understanding the Hardware Limits of Sparse Neural Networks30 Sep 2023 0 repositories listed
-
A rigorous benchmarking of methods for SARS-CoV-2 lineage abundance estimation in wastewater29 Sep 2023 0 repositories listed
-
Benchmarking and In-depth Performance Study of Large Language Models on Habana Gaudi Processors29 Sep 2023 0 repositories listed
-
Benchmarking Collaborative Learning Methods Cost-Effectiveness for Prostate Segmentation29 Sep 2023 0 repositories listed
-
Intuitive or Dependent? Investigating LLMs' Behavior Style to Conflicting Prompts29 Sep 2023 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.