Browse State-of-the-Art › Benchmarking › Papers, page 39
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 39 of 56: papers 3,801 to 3,900 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
CubeSat-Enabled Free-Space Optics: Joint Data Communication and Fine Beam Tracking13 Jun 2024 0 repositories listed
-
Decoding the Diversity: A Review of the Indic AI Research Landscape13 Jun 2024 0 repositories listed
-
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living13 Jun 2024 0 repositories listed
-
ResearchArena: Benchmarking LLMs' Ability to Collect and Organize Information as Research Agents13 Jun 2024 0 repositories listed
-
How well it works: Benchmarking performance of GPT models on medical natural language processing tasks12 Jun 2024 0 repositories listed
-
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives12 Jun 2024 0 repositories listed
-
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets12 Jun 2024 0 repositories listed
-
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents12 Jun 2024 0 repositories listed
-
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases12 Jun 2024 0 repositories listed
-
A PRISMA Driven Systematic Review of Publicly Available Datasets for Benchmark and Model Developments for Industrial Defect Detection11 Jun 2024 0 repositories listed
-
Advancing Annotation of Stance in Social Media Posts: A Comparative Analysis of Large Language Models and Crowd Sourcing11 Jun 2024 0 repositories listed
-
Benchmarking and Boosting Radiology Report Generation for 3D High-Resolution Medical Images11 Jun 2024 0 repositories listed
-
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models11 Jun 2024 0 repositories listed
-
DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition11 Jun 2024 0 repositories listed
-
Can Language Models Serve as Text-Based World Simulators?10 Jun 2024 0 repositories listed
-
Data-driven Power Flow Linearization: Simulation10 Jun 2024 0 repositories listed
-
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking10 Jun 2024 0 repositories listed
-
1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation8 Jun 2024 0 repositories listed
-
Benchmarking Neural Decoding Backbones towards Enhanced On-edge iBCI Applications8 Jun 2024 0 repositories listed
-
Behavior Structformer: Learning Players Representations with Structured Tokenization7 Jun 2024 0 repositories listed
-
GenzIQA: Generalized Image Quality Assessment using Prompt-Guided Latent Diffusion Models7 Jun 2024 0 repositories listed
-
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation7 Jun 2024 0 repositories listed
-
Scenarios and Approaches for Situated Natural Language Explanations7 Jun 2024 0 repositories listed
-
6 Jun 2024 0 repositories listed
-
Benchmarking AlphaFold3's protein-protein complex accuracy and machine learning prediction reliability for binding free energy changes upon mutation6 Jun 2024 0 repositories listed
-
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices6 Jun 2024 0 repositories listed
-
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning6 Jun 2024 0 repositories listed
-
Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking6 Jun 2024 0 repositories listed
-
Performance of large language models in numerical vs. semantic medical knowledge: Benchmarking on evidence-based Q&As6 Jun 2024 0 repositories listed
-
Statistical Multicriteria Benchmarking via the GSD-Front6 Jun 2024 0 repositories listed
-
Comparative Benchmarking of Failure Detection Methods in Medical Image Segmentation: Unveiling the Role of Confidence Aggregation5 Jun 2024 0 repositories listed
-
Bi-DCSpell: A Bi-directional Detector-Corrector Interactive Framework for Chinese Spelling Check4 Jun 2024 0 repositories listed
-
Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs4 Jun 2024 0 repositories listed
-
ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection3 Jun 2024 0 repositories listed
-
LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions3 Jun 2024 0 repositories listed
-
R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models3 Jun 2024 0 repositories listed
-
Scaffold Splits Overestimate Virtual Screening Performance2 Jun 2024 0 repositories listed
-
On the project risk baseline: integrating aleatory uncertainty into project scheduling31 May 2024 0 repositories listed
-
Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images30 May 2024 0 repositories listed
-
Categorization of 33 computational methods to detect spatially variable genes from spatially resolved transcriptomics data29 May 2024 0 repositories listed
-
MDIW-13: a New Multi-Lingual and Multi-Script Database and Benchmark for Script Identification29 May 2024 0 repositories listed
-
Exploring Thermography Technology: A Comprehensive Facial Dataset for Face Detection, Recognition, and Emotion28 May 2024 0 repositories listed
-
Risk-Neutral Generative Networks28 May 2024 0 repositories listed
-
A Correlation- and Mean-Aware Loss Function and Benchmarking Framework to Improve GAN-based Tabular Data Synthesis27 May 2024 0 repositories listed
-
Benchmarking General-Purpose In-Context Learning27 May 2024 0 repositories listed
-
BOLD: Boolean Logic Deep Learning25 May 2024 0 repositories listed
-
GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases25 May 2024 0 repositories listed
-
Application based Evaluation of an Efficient Spike-Encoder, "Spiketrum"24 May 2024 0 repositories listed
-
Benchmarking Hierarchical Image Pyramid Transformer for the classification of colon biopsies and polyps in histopathology images24 May 2024 0 repositories listed
-
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks24 May 2024 0 repositories listed
-
Free Performance Gain from Mixing Multiple Partially Labeled Samples in Multi-label Image Classification24 May 2024 0 repositories listed
-
Full-stack evaluation of Machine Learning inference workloads for RISC-V systems24 May 2024 0 repositories listed
-
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study24 May 2024 0 repositories listed
-
MCDFN: Supply Chain Demand Forecasting via an Explainable Multi-Channel Data Fusion Network Model24 May 2024 0 repositories listed
-
A Gap in Time: The Challenge of Processing Heterogeneous IoT Data in Digitalized Buildings23 May 2024 0 repositories listed
-
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models22 May 2024 0 repositories listed
-
CT-Eval: Benchmarking Chinese Text-to-Table Performance in Large Language Models20 May 2024 0 repositories listed
-
EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods20 May 2024 0 repositories listed
-
EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models18 May 2024 0 repositories listed
-
BraTS-Path Challenge: Assessing Heterogeneous Histopathologic Brain Tumor Sub-regions17 May 2024 0 repositories listed
-
From Generalist to Specialist: Improving Large Language Models for Medical Physics Using ARCoT17 May 2024 0 repositories listed
-
SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge17 May 2024 0 repositories listed
-
A Robust Autoencoder Ensemble-Based Approach for Anomaly Detection in Text16 May 2024 0 repositories listed
-
SpeechVerse: A Large-scale Generalizable Audio Language Model14 May 2024 0 repositories listed
-
Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness13 May 2024 0 repositories listed
-
Comparative analysis of neural network architectures for short-term FOREX forecasting13 May 2024 0 repositories listed
-
oTTC: Object Time-to-Contact for Motion Estimation in Autonomous Driving13 May 2024 0 repositories listed
-
UCCIX: Irish-eXcellence Large Language Model13 May 2024 0 repositories listed
-
Benchmarking Cross-Domain Audio-Visual Deception Detection11 May 2024 0 repositories listed
-
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs10 May 2024 0 repositories listed
-
Agent-oriented Joint Decision Support for Data Owners in Auction-based Federated Learning9 May 2024 0 repositories listed
-
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking7 May 2024 0 repositories listed
-
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images6 May 2024 0 repositories listed
-
ATG: Benchmarking Automated Theorem Generation for Generative Language Models5 May 2024 0 repositories listed
-
Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models5 May 2024 0 repositories listed
-
PhilHumans: Benchmarking Machine Learning for Personal Health4 May 2024 0 repositories listed
-
Systematic Review: Anomaly Detection in Connected and Autonomous Vehicles4 May 2024 0 repositories listed
-
A Normative Framework for Benchmarking Consumer Fairness in Large Language Model Recommender System3 May 2024 0 repositories listed
-
Toward end-to-end interpretable convolutional neural networks for waveform signals3 May 2024 0 repositories listed
-
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News2 May 2024 0 repositories listed
-
Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attribution Methods2 May 2024 0 repositories listed
-
CityLearn v2: Energy-flexible, resilient, occupant-centric, and carbon-aware management of grid-interactive communities2 May 2024 0 repositories listed
-
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting30 Apr 2024 0 repositories listed
-
Evaluating Deep Clustering Algorithms on Non-Categorical 3D CAD Models29 Apr 2024 0 repositories listed
-
MileBench: Benchmarking MLLMs in Long Context29 Apr 2024 0 repositories listed
-
On the Impact of Data Heterogeneity in Federated Learning Environments with Application to Healthcare Networks29 Apr 2024 0 repositories listed
-
Efficient Exploration of Image Classifier Failures with Bayesian Optimization and Text-to-Image Models26 Apr 2024 0 repositories listed
-
Stochastic Spiking Neural Networks with First-to-Spike Coding26 Apr 2024 0 repositories listed
-
Benchmarking Mobile Device Control Agents across Diverse Configurations25 Apr 2024 0 repositories listed
-
24 Apr 2024 0 repositories listed Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Empirical Analysis of the Dynamic Binary Value Problem with IOHprofiler24 Apr 2024 0 repositories listed
-
Benchmarking Advanced Text Anonymisation Methods: A Comparative Study on Novel and Traditional Approaches22 Apr 2024 0 repositories listed
-
EnzChemRED, a rich enzyme chemistry relation extraction dataset22 Apr 2024 0 repositories listed
-
Open Datasets for Satellite Radio Resource Control22 Apr 2024 0 repositories listed
-
TeamTrack: A Dataset for Multi-Sport Multi-Object Tracking in Full-pitch Videos22 Apr 2024 0 repositories listed
-
The Adversarial AI-Art: Understanding, Generation, Detection, and Benchmarking22 Apr 2024 0 repositories listed
-
Authentic Emotion Mapping: Benchmarking Facial Expressions in Real News21 Apr 2024 0 repositories listed
-
In-situ process monitoring and adaptive quality enhancement in laser additive manufacturing: a critical review21 Apr 2024 0 repositories listed
-
Bridging the Gap Between Theory and Practice: Benchmarking Transfer Evolutionary Optimization20 Apr 2024 0 repositories listed
-
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning19 Apr 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.