Browse State-of-the-Art › Benchmarking › Papers, page 40
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 40 of 56: papers 3,901 to 4,000 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Integrated Sensing and Communication enabled Multiple Base Stations Cooperative UAV Detection19 Apr 2024 0 repositories listed
-
Environment-aware UAV Communications: CKM Construction and Predictive Beamforming18 Apr 2024 0 repositories listed
-
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions17 Apr 2024 0 repositories listed
-
Neural Network Approach for Non-Markovian Dissipative Dynamics of Many-Body Open Quantum Systems17 Apr 2024 0 repositories listed
-
Benchmarking changepoint detection algorithms on cardiac time series16 Apr 2024 0 repositories listed
-
Data Collection of Real-Life Knowledge Work in Context: The RLKWiC Dataset16 Apr 2024 0 repositories listed
-
Iterated Invariant Extended Kalman Filter (IterIEKF)16 Apr 2024 0 repositories listed
-
Neuromorphic Vision-based Motion Segmentation with Graph Transformer Neural Network16 Apr 2024 0 repositories listed
-
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs16 Apr 2024 0 repositories listed
-
A Universal Protocol to Benchmark Camera Calibration for Sports15 Apr 2024 0 repositories listed
-
Feature selection in linear SVMs via a hard cardinality constraint: a scalable SDP decomposition approach15 Apr 2024 0 repositories listed
-
LLM Evaluators Recognize and Favor Their Own Generations15 Apr 2024 0 repositories listed
-
MMInA: Benchmarking Multihop Multimodal Internet Agents15 Apr 2024 0 repositories listed
-
Practical Guidelines for Cell Segmentation Models Under Optical Aberrations in Microscopy12 Apr 2024 0 repositories listed
-
Exploring the Decentraland Economy: Multifaceted Parcel Attributes, Key Insights, and Benchmarking11 Apr 2024 0 repositories listed
-
Certifying almost all quantum states with few single-qubit measurements10 Apr 2024 0 repositories listed
-
10 Apr 2024 0 repositories listed
-
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution9 Apr 2024 0 repositories listed
-
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs9 Apr 2024 0 repositories listed
-
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering8 Apr 2024 0 repositories listed
-
A Comparison of Cryptocurrency Volatility-benchmarking New and Mature Asset Classes7 Apr 2024 0 repositories listed
-
Multicalibration for Confidence Scoring in LLMs6 Apr 2024 0 repositories listed
-
SDFR: Synthetic Data for Face Recognition Competition6 Apr 2024 0 repositories listed
-
Dynamic Risk Assessment Methodology with an LDM-based System for Parking Scenarios5 Apr 2024 0 repositories listed
-
GNNBENCH: Fair and Productive Benchmarking for Single-GPU GNN System5 Apr 2024 0 repositories listed
-
DiffBody: Human Body Restoration by Imagining with Generative Diffusion Prior4 Apr 2024 0 repositories listed
-
NL2KQL: From Natural Language to Kusto Query3 Apr 2024 0 repositories listed
-
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach2 Apr 2024 0 repositories listed
-
On the reduction of Linear Parameter-Varying State-Space models2 Apr 2024 0 repositories listed
-
Diffusion-Driven Domain Adaptation for Generating 3D Molecules1 Apr 2024 0 repositories listed
-
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations1 Apr 2024 0 repositories listed
-
SpiralMLP: A Lightweight Vision MLP Architecture31 Mar 2024 0 repositories listed
-
Comparing Hyper-optimized Machine Learning Models for Predicting Efficiency Degradation in Organic Solar Cells29 Mar 2024 0 repositories listed
-
Benchmarking Image Transformers for Prostate Cancer Detection from Ultrasound Data27 Mar 2024 0 repositories listed
-
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination26 Mar 2024 0 repositories listed
-
Benchmarking Video Frame Interpolation25 Mar 2024 0 repositories listed
-
25 Mar 2024 0 repositories listed
-
TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring23 Mar 2024 0 repositories listed
-
Broadening the Scope of Neural Network Potentials through Direct Inclusion of Additional Molecular Attributes22 Mar 2024 0 repositories listed
-
Unifying Large Language Model and Deep Reinforcement Learning for Human-in-Loop Interactive Socially-aware Navigation22 Mar 2024 0 repositories listed
-
Subjective Quality Assessment of Compressed Tone-Mapped High Dynamic Range Videos22 Mar 2024 0 repositories listed
-
Transactive Local Energy Markets Enable Community-Level Resource Coordination Using Individual Rewards22 Mar 2024 0 repositories listed
-
ChatGPT Alternative Solutions: Large Language Models Survey21 Mar 2024 0 repositories listed
-
Embarrassingly Simple Scribble Supervision for 3D Medical Segmentation19 Mar 2024 0 repositories listed
-
Benchmarking Badminton Action Recognition with a New Fine-Grained Dataset19 Mar 2024 0 repositories listed
-
A Sober Look at the Robustness of CLIPs to Spurious Features18 Mar 2024 0 repositories listed
-
Leveraging Spatial and Semantic Feature Extraction for Skin Cancer Diagnosis with Capsule Networks and Graph Neural Networks18 Mar 2024 0 repositories listed
-
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety18 Mar 2024 0 repositories listed
-
FlowMind: Automatic Workflow Generation with LLMs17 Mar 2024 0 repositories listed
-
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking17 Mar 2024 0 repositories listed
-
Depression Detection on Social Media with Large Language Models16 Mar 2024 0 repositories listed
-
Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-adaptive Attacks15 Mar 2024 0 repositories listed
-
A tutorial on multi-view autoencoders using the multi-view-AE library12 Mar 2024 0 repositories listed
-
IndicSTR12: A Dataset for Indic Scene Text Recognition12 Mar 2024 0 repositories listed
-
An Approach to Evaluate Modeling Adequacy for Small-Signal Stability Analysis of IBR-related SSOs in Multimachine Systems12 Mar 2024 0 repositories listed
-
A Holistic Framework Towards Vision-based Traffic Signal Control with Microscopic Simulation11 Mar 2024 0 repositories listed
-
(N,K)-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model11 Mar 2024 0 repositories listed
-
Exploring the Adversarial Frontier: Quantifying Robustness via Adversarial Hypervolume8 Mar 2024 0 repositories listed
-
Benchmarking News Recommendation in the Era of Green AI7 Mar 2024 0 repositories listed
-
NLPre: a revised approach towards language-centric benchmarking of Natural Language Preprocessing systems7 Mar 2024 0 repositories listed
-
A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Video6 Mar 2024 0 repositories listed
-
6 Mar 2024 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking the Text-to-SQL Capability of Large Language Models: A Comprehensive Evaluation5 Mar 2024 0 repositories listed
-
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering5 Mar 2024 0 repositories listed
-
Classification of the Fashion-MNIST Dataset on a Quantum Computer4 Mar 2024 0 repositories listed
-
Model Lakes4 Mar 2024 0 repositories listed
-
Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground4 Mar 2024 0 repositories listed
-
A Bayesian Committee Machine Potential for Oxygen-containing Organic Compounds2 Mar 2024 0 repositories listed
-
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance1 Mar 2024 0 repositories listed
-
Beyond Single-Model Views for Deep Learning: Optimization versus Generalizability of Stochastic Optimization Algorithms1 Mar 2024 0 repositories listed
-
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models1 Mar 2024 0 repositories listed
-
SINDy vs Hard Nonlinearities and Hidden Dynamics: a Benchmarking Study1 Mar 2024 0 repositories listed
-
The 6th Affective Behavior Analysis in-the-wild (ABAW) Competition29 Feb 2024 0 repositories listed
-
A Large-scale Evaluation of Pretraining Paradigms for the Detection of Defects in Electroluminescence Solar Cell Images27 Feb 2024 0 repositories listed
-
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies27 Feb 2024 0 repositories listed
-
The Seeker's Dilemma: Realistic Formulation and Benchmarking for Hardware Trojan Detection27 Feb 2024 0 repositories listed
-
Benchmarking LLMs on the Semantic Overlap Summarization Task26 Feb 2024 0 repositories listed
-
Performance Comparison of Surrogate-Assisted Evolutionary Algorithms on Computational Fluid Dynamics Problems26 Feb 2024 0 repositories listed
-
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset26 Feb 2024 0 repositories listed
-
E(3)-equivariant models cannot learn chirality: Field-based molecular generation24 Feb 2024 0 repositories listed
-
Decoding Intelligence: A Framework for Certifying Knowledge Comprehension in LLMs24 Feb 2024 0 repositories listed
-
Benchmarking Observational Studies with Experimental Data under Right-Censoring23 Feb 2024 0 repositories listed
-
Benchmarking the Robustness of Panoptic Segmentation for Automated Driving23 Feb 2024 0 repositories listed
-
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models21 Feb 2024 0 repositories listed
-
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models21 Feb 2024 0 repositories listed
-
KetGPT -- Dataset Augmentation of Quantum Circuits using Transformers20 Feb 2024 0 repositories listed
-
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation19 Feb 2024 0 repositories listed
-
Learning Disentangled Audio Representations through Controlled Synthesis16 Feb 2024 0 repositories listed
-
VATr++: Choose Your Words Wisely for Handwritten Text Generation16 Feb 2024 0 repositories listed
-
Benchmarking federated strategies in Peer-to-Peer Federated learning for biomedical data15 Feb 2024 0 repositories listed
-
Large-scale Benchmarking of Metaphor-based Optimization Heuristics15 Feb 2024 0 repositories listed
-
Multi-Fidelity Methods for Optimization: A Survey15 Feb 2024 0 repositories listed
-
Recommendations for Baselines and Benchmarking Approximate Gaussian Processes15 Feb 2024 0 repositories listed
-
Design and Realization of a Benchmarking Testbed for Evaluating Autonomous Platooning Algorithms14 Feb 2024 0 repositories listed
-
Evaluation of simulation methods for tumor subclonal reconstruction14 Feb 2024 0 repositories listed
-
Privacy-Preserving Language Model Inference with Instance Obfuscation13 Feb 2024 0 repositories listed
-
Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT12 Feb 2024 0 repositories listed
-
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages12 Feb 2024 0 repositories listed
-
Impact of spatial transformations on landscape features of CEC2022 basic benchmark problems12 Feb 2024 0 repositories listed
-
Estimating the Effect of Crosstalk Error on Circuit Fidelity Using Noisy Intermediate-Scale Quantum Devices10 Feb 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.