Browse State-of-the-Art › Benchmarking › Papers, page 44
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 44 of 56: papers 4,301 to 4,400 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Enhancing Navigation Benchmarking and Perception Data Generation for Row-based Crops in Simulation27 Jun 2023 0 repositories listed
-
Paradigm Shift in Sustainability Disclosure Analysis: Empowering Stakeholders with CHATREPORT, a Language Model-Based Tool27 Jun 2023 0 repositories listed
-
Pulse Shape-Aided Multipath Delay Estimation for Fine-Grained WiFi Sensing27 Jun 2023 0 repositories listed
-
Hybrid Precoder and Combiner Designs for Decentralized Parameter Estimation in mmWave MIMO Wireless Sensor Networks25 Jun 2023 0 repositories listed
-
Improving Reference-based Distinctive Image Captioning with Contrastive Rewards25 Jun 2023 0 repositories listed
-
A Comprehensive Study on the Robustness of Image Classification and Object Detection in Remote Sensing: Surveying and Benchmarking21 Jun 2023 0 repositories listed
-
Evaluation of Popular XAI Applied to Clinical Prediction Models: Can They be Trusted?21 Jun 2023 0 repositories listed
-
On Evaluation of Document Classification using RVL-CDIP21 Jun 2023 0 repositories listed
-
Diverse Community Data for Benchmarking Data Privacy Algorithms20 Jun 2023 0 repositories listed
-
Benchmarking Robustness of Deep Reinforcement Learning approaches to Online Portfolio Management19 Jun 2023 0 repositories listed
-
Fairness Index Measures to Evaluate Bias in Biometric Recognition19 Jun 2023 0 repositories listed
-
Formal Covariate Benchmarking to Bound Omitted Variable Bias18 Jun 2023 0 repositories listed
-
MA-BBOB: Many-Affine Combinations of BBOB Functions for Evaluating AutoML Approaches in Noiseless Numerical Black-Box Optimization Contexts18 Jun 2023 0 repositories listed
-
Benchmarking Deep Learning Architectures for Urban Vegetation Point Cloud Semantic Segmentation from MLS17 Jun 2023 0 repositories listed
-
ALP: Action-Aware Embodied Learning for Perception16 Jun 2023 0 repositories listed
-
DiPlomat: A Dialogue Dataset for Situated Pragmatic Reasoning15 Jun 2023 0 repositories listed
-
DISC: a Dataset for Integrated Sensing and Communication in mmWave Systems15 Jun 2023 0 repositories listed
-
Large-Scale Quantum Separability Through a Reproducible Machine Learning Lens15 Jun 2023 0 repositories listed
-
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion15 Jun 2023 0 repositories listed
-
RRSIS: Referring Remote Sensing Image Segmentation14 Jun 2023 0 repositories listed
-
A Cloud-based Machine Learning Pipeline for the Efficient Extraction of Insights from Customer Reviews13 Jun 2023 0 repositories listed
-
Contribution à l'Optimisation d'un Comportement Collectif pour un Groupe de Robots Autonomes10 Jun 2023 0 repositories listed
-
9 Jun 2023 0 repositories listed
-
Share, Collaborate, Benchmark: Advancing Travel Demand Research through rigorous open-source collaboration9 Jun 2023 0 repositories listed
-
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems8 Jun 2023 0 repositories listed
-
Benchmarking Foundation Models with Language-Model-as-an-Examiner7 Jun 2023 0 repositories listed
-
ICON²: Reliably Benchmarking Predictive Inequity in Object Detection7 Jun 2023 0 repositories listed
-
Improved statistical benchmarking of digital pathology models using pairwise frames evaluation7 Jun 2023 0 repositories listed
-
RD-Suite: A Benchmark for Ranking Distillation7 Jun 2023 0 repositories listed
-
Applying Standards to Advance Upstream & Downstream Ethics in Large Language Models6 Jun 2023 0 repositories listed
-
Benchmarking Robustness of AI-Enabled Multi-sensor Fusion Systems: Challenges and Opportunities6 Jun 2023 0 repositories listed
-
Explainable AI using expressive Boolean formulas6 Jun 2023 0 repositories listed
-
Financial Numeric Extreme Labelling: A Dataset and Benchmarking for XBRL Tagging6 Jun 2023 0 repositories listed
-
N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition5 Jun 2023 0 repositories listed
-
EfficientSRFace: An Efficient Network with Super-Resolution Enhancement for Accurate Face Detection4 Jun 2023 0 repositories listed
-
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning4 Jun 2023 0 repositories listed
-
3 Jun 2023 0 repositories listed
-
Benchmarking Robustness of Adaptation Methods on Pre-trained Vision-Language Models3 Jun 2023 0 repositories listed
-
Break a Lag: Triple Exponential Moving Average for Enhanced Optimization2 Jun 2023 0 repositories listed
-
Hybrid Long Document Summarization using C2F-FAR and ChatGPT: A Practical Study1 Jun 2023 0 repositories listed
-
HySpecNet-11k: A Large-Scale Hyperspectral Dataset for Benchmarking Learning-Based Hyperspectral Image Compression Methods1 Jun 2023 0 repositories listed
-
The Brain Tumor Segmentation (BraTS-METS) Challenge 2023: Brain Metastasis Segmentation on Pre-treatment MRI1 Jun 2023 0 repositories listed
-
1 Jun 2023 0 repositories listed
-
Human Body Shape Classification Based on a Single Image29 May 2023 0 repositories listed
-
Benchmarking Diverse-Modal Entity Linking with Generative Models27 May 2023 0 repositories listed
-
Exploring the Practicality of Generative Retrieval on Dynamic Corpora27 May 2023 0 repositories listed
-
Benchmarking state-of-the-art gradient boosting algorithms for classification26 May 2023 0 repositories listed
-
Analysis of modular CMA-ES on strict box-constrained problems in the SBOX-COST benchmarking suite24 May 2023 0 repositories listed
-
Barkour: Benchmarking Animal-level Agility with Quadruped Robots24 May 2023 0 repositories listed
-
LAraBench: Benchmarking Arabic AI with Large Language Models24 May 2023 0 repositories listed
-
BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer24 May 2023 0 repositories listed
-
Domain-Expanded ASTE: Rethinking Generalization in Aspect Sentiment Triplet Extraction23 May 2023 0 repositories listed
-
Multilingual Large Language Models Are Not (Yet) Code-Switchers23 May 2023 0 repositories listed
-
R2H: Building Multimodal Navigation Helpers that Respond to Help Requests23 May 2023 0 repositories listed
-
Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via Debate22 May 2023 0 repositories listed
-
Value-at-Risk-Based Portfolio Insurance: Performance Evaluation and Benchmarking Against CPPI in a Markov-Modulated Regime-Switching Market21 May 2023 0 repositories listed
-
Patterns of Convergence and Bound Constraint Violation in Differential Evolution on SBOX-COST Benchmarking Suite20 May 2023 0 repositories listed
-
TELeR: A General Taxonomy of LLM Prompts for Benchmarking Complex Tasks19 May 2023 0 repositories listed
-
Ahead-of-Time P-Tuning18 May 2023 0 repositories listed
-
Benchmarking Deep Learning Frameworks for Automated Diagnosis of Ocular Toxoplasmosis: A Comprehensive Approach to Classification and Segmentation18 May 2023 0 repositories listed
-
Boost Vision Transformer with GPU-Friendly Sparsity and Quantization18 May 2023 0 repositories listed
-
Human Behavioral Benchmarking: Numeric Magnitude Comparison Effects in Large Language Models18 May 2023 0 repositories listed
-
17 May 2023 0 repositories listed
-
Smiling Women Pitching Down: Auditing Representational and Presentational Gender Biases in Image Generative AI17 May 2023 0 repositories listed
-
Towards More Robust NLP System Evaluation: Handling Missing Scores in Benchmarks17 May 2023 0 repositories listed
-
DLUE: Benchmarking Document Language Understanding16 May 2023 0 repositories listed
-
Benchmarking the human brain against computational architectures15 May 2023 0 repositories listed
-
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking15 May 2023 0 repositories listed
-
Predictive Models from Quantum Computer Benchmarks15 May 2023 0 repositories listed
-
A Strong Sustainability Paradigm Based Analytical Hierarchy Process (SSP-AHP) Method to Evaluate Sustainable Healthcare Systems13 May 2023 0 repositories listed
-
MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine12 May 2023 0 repositories listed
-
Uncertainty in GNN Learning Evaluations: The Importance of a Consistent Benchmark for Community Detection10 May 2023 0 repositories listed
-
Comparing Foundation Models using Data Kernels9 May 2023 0 repositories listed
-
A Comprehensive Study on Dataset Distillation: Performance, Privacy, Robustness and Fairness5 May 2023 0 repositories listed
-
Towards Segment Anything Model (SAM) for Medical Image Segmentation: A Survey5 May 2023 0 repositories listed
-
Semantic Segmentation using Vision Transformers: A survey5 May 2023 0 repositories listed
-
Analyzing Hong Kong's Legal Judgments from a Computational Linguistics point-of-view4 May 2023 0 repositories listed
-
Can LLMs Capture Human Preferences?4 May 2023 0 repositories listed
-
A Simulation-Augmented Benchmarking Framework for Automatic RSO Streak Detection in Single-Frame Space Images30 Apr 2023 0 repositories listed
-
Benchmarking Automated Machine Learning Methods for Price Forecasting Applications28 Apr 2023 0 repositories listed
-
ChatGPT vs State-of-the-Art Models: A Benchmarking Study in Keyphrase Generation Task27 Apr 2023 0 repositories listed
-
Scalable, Distributed AI Frameworks: Leveraging Cloud Computing for Enhanced Deep Learning Performance and Efficiency26 Apr 2023 0 repositories listed
-
CIMLA: Interpretable AI for inference of differential causal networks25 Apr 2023 0 repositories listed
-
Unsupervised Synthetic Image Refinement via Contrastive Learning and Consistent Semantic-Structural Constraints25 Apr 2023 0 repositories listed
-
A Framework for Benchmarking Real-Time Embedded Object Detection23 Apr 2023 0 repositories listed
-
Vision Transformer for Efficient Chest X-ray and Gastrointestinal Image Classification23 Apr 2023 0 repositories listed
-
Learning a quantum computer's capability20 Apr 2023 0 repositories listed
-
Towards a Benchmark for Scientific Understanding in Humans and Machines20 Apr 2023 0 repositories listed
-
Computational and Exploratory Landscape Analysis of the GKLS Generator18 Apr 2023 0 repositories listed
-
UDTIRI: An Online Open-Source Intelligent Road Inspection Benchmark Suite18 Apr 2023 0 repositories listed
-
OOD-CV-v2: An extended Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images17 Apr 2023 0 repositories listed
-
Dialogue Games for Benchmarking Language Understanding: Motivation, Taxonomy, Strategy14 Apr 2023 0 repositories listed
-
Benchmarking the Physical-world Adversarial Robustness of Vehicle Detection11 Apr 2023 0 repositories listed
-
Improving Items and Contexts Understanding with Descriptive Graph for Conversational Recommendation11 Apr 2023 0 repositories listed
-
10 Apr 2023 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
ForamViT-GAN: Exploring New Paradigms in Deep Learning for Micropaleontological Image Analysis9 Apr 2023 0 repositories listed
-
Benchmarking the Robustness of Quantized Models8 Apr 2023 0 repositories listed
-
DRAC: Diabetic Retinopathy Analysis Challenge with Ultra-Wide Optical Coherence Tomography Angiography Images5 Apr 2023 0 repositories listed
-
OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI4 Apr 2023 0 repositories listed
-
A Latent Fingerprint in the Wild Database3 Apr 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.