Browse State-of-the-Art › Benchmarking › Papers, page 30
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 30 of 56: papers 2,901 to 3,000 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
JointDistill: Adaptive Multi-Task Distillation for Joint Depth Estimation and Scene Segmentation15 May 2025 0 repositories listed
-
Real-World fNIRS-Based Brain-Computer Interfaces: Benchmarking Deep Learning and Classical Models in Interactive Gaming15 May 2025 0 repositories listed
-
Visual Fidelity Index for Generative Semantic Communications with Critical Information Embedding15 May 2025 0 repositories listed
-
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs15 May 2025 0 repositories listed
-
A Standardized Benchmark Set of Clustering Problem Instances for Comparing Black-Box Optimizers14 May 2025 0 repositories listed
-
How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference14 May 2025 0 repositories listed
-
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning14 May 2025 0 repositories listed
-
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation14 May 2025 0 repositories listed
-
RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo14 May 2025 0 repositories listed
-
TARGET: Benchmarking Table Retrieval for Generative Tasks14 May 2025 0 repositories listed
-
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts14 May 2025 0 repositories listed
-
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models14 May 2025 0 repositories listed
-
A Large-scale Benchmark on Geological Fault Delineation Models: Domain Shift, Training Dynamics, Generalizability, Evaluation and Inferential Behavior13 May 2025 0 repositories listed
-
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities13 May 2025 0 repositories listed
-
Load-independent Metrics for Benchmarking Force Controllers13 May 2025 0 repositories listed
-
Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 203012 May 2025 0 repositories listed
-
Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs12 May 2025 0 repositories listed
-
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems12 May 2025 0 repositories listed
-
Benchmarking Retrieval-Augmented Generation for Chemistry12 May 2025 0 repositories listed
-
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning12 May 2025 0 repositories listed
-
PRISM: Complete Online Decentralized Multi-Agent Pathfinding with Rapid Information Sharing using Motion Constraints12 May 2025 0 repositories listed
-
The Pitfalls of Benchmarking in Algorithm Selection: What We Are Getting Wrong12 May 2025 0 repositories listed
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration11 May 2025 0 repositories listed
-
Optimizing Recommendations using Fine-Tuned LLMs11 May 2025 0 repositories listed
-
Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA9 May 2025 0 repositories listed
-
Evaluating Financial Sentiment Analysis with Annotators Instruction Assisted Prompting: Enhancing Contextual Interpretation and Stock Prediction Accuracy9 May 2025 0 repositories listed
-
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information9 May 2025 0 repositories listed
-
Autoregressive Stochastic Clock Jitter Compensation in Analog-to-Digital Converters8 May 2025 0 repositories listed
-
Benchmarking Ophthalmology Foundation Models for Clinically Significant Age Macular Degeneration Detection8 May 2025 0 repositories listed
-
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations8 May 2025 0 repositories listed
-
Federated Deconfounding and Debiasing Learning for Out-of-Distribution Generalization8 May 2025 0 repositories listed
-
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation8 May 2025 0 repositories listed
-
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents8 May 2025 0 repositories listed
-
Alpha Excel Benchmark7 May 2025 0 repositories listed
-
Call for Action: towards the next generation of symbolic regression benchmark6 May 2025 0 repositories listed
-
Completing Spatial Transcriptomics Data for Gene Expression Prediction Benchmarking5 May 2025 0 repositories listed
-
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities5 May 2025 0 repositories listed
-
Representation Learning of Limit Order Book: A Comprehensive Study and Benchmarking4 May 2025 0 repositories listed
-
BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models3 May 2025 0 repositories listed
-
CMAWRNet: Multiple Adverse Weather Removal via a Unified Quaternion Neural Architecture3 May 2025 0 repositories listed
-
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey3 May 2025 0 repositories listed
-
Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking3 May 2025 0 repositories listed
-
Not Every Tree Is a Forest: Benchmarking Forest Types from Satellite Remote Sensing3 May 2025 0 repositories listed
-
PhytoSynth: Leveraging Multi-modal Generative Models for Crop Disease Data Generation with Novel Benchmarking and Prompt Engineering Approach3 May 2025 0 repositories listed
-
Can Foundation Models Really Segment Tumors? A Benchmarking Odyssey in Lung CT Imaging2 May 2025 0 repositories listed
-
Overview and practical recommendations on using Shapley Values for identifying predictive biomarkers via CATE modeling2 May 2025 0 repositories listed
-
AI-ready Snow Radar Echogram Dataset (SRED) for climate change monitoring1 May 2025 0 repositories listed
-
EnronQA: Towards Personalized RAG over Private Documents1 May 2025 0 repositories listed
-
InterLoc: LiDAR-based Intersection Localization using Road Segmentation with Automated Evaluation Method1 May 2025 0 repositories listed
-
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation1 May 2025 0 repositories listed
-
From Precision to Perception: User-Centred Evaluation of Keyword Extraction Algorithms for Internet-Scale Contextual Advertising30 Apr 2025 0 repositories listed
-
Sadeed: Advancing Arabic Diacritization Through Small Language Model30 Apr 2025 0 repositories listed
-
Towards Robust and Generalizable Gerchberg Saxton based Physics Inspired Neural Networks for Computer Generated Holography: A Sensitivity Analysis Framework30 Apr 2025 0 repositories listed
-
Evaluating Generative Models for Tabular Data: Novel Metrics and Benchmarking29 Apr 2025 0 repositories listed
-
Hydra: Marker-Free RGB-D Hand-Eye Calibration29 Apr 2025 0 repositories listed
-
LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs29 Apr 2025 0 repositories listed
-
On the Potential of Large Language Models to Solve Semantics-Aware Process Mining Tasks29 Apr 2025 0 repositories listed
-
SecRepoBench: Benchmarking LLMs for Secure Code Generation in Real-World Repositories29 Apr 2025 0 repositories listed
-
The Leaderboard Illusion29 Apr 2025 0 repositories listed
-
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets28 Apr 2025 0 repositories listed
-
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies28 Apr 2025 0 repositories listed
-
WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution28 Apr 2025 0 repositories listed
-
Quantitative evaluation of brain-inspired vision sensors in high-speed robotic perception27 Apr 2025 0 repositories listed
-
The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach27 Apr 2025 0 repositories listed
-
Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis25 Apr 2025 0 repositories listed
-
Design and benchmarking of a two degree of freedom tendon driver unit for cable-driven wearable technologies24 Apr 2025 0 repositories listed
-
QuantBench: Benchmarking AI Methods for Quantitative Investment24 Apr 2025 0 repositories listed
-
Token Sequence Compression for Efficient Multimodal Computing24 Apr 2025 0 repositories listed
-
From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories23 Apr 2025 0 repositories listed
-
A Large-scale Class-level Benchmark Dataset for Code Generation with LLMs22 Apr 2025 0 repositories listed
-
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V322 Apr 2025 0 repositories listed
-
Benchmarking machine learning models for predicting aerofoil performance22 Apr 2025 0 repositories listed
-
CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents22 Apr 2025 0 repositories listed
-
Enhancing TCR-Peptide Interaction Prediction with Pretrained Language Models and Molecular Representations22 Apr 2025 0 repositories listed
-
Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room22 Apr 2025 0 repositories listed
-
Audio-Visual Class-Incremental Learning for Fish Feeding intensity Assessment in Aquaculture21 Apr 2025 0 repositories listed
-
Establishing Reliability Metrics for Reward Models in Large Language Models21 Apr 2025 0 repositories listed
-
Speaker Fuzzy Fingerprints: Benchmarking Text-Based Identification in Multiparty Dialogues21 Apr 2025 0 repositories listed
-
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents20 Apr 2025 0 repositories listed
-
IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays20 Apr 2025 0 repositories listed
-
AI Idea Bench 2025: AI Research Idea Generation Benchmark19 Apr 2025 0 repositories listed
-
Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation19 Apr 2025 0 repositories listed
-
CodeCrash: Stress Testing LLM Reasoning under Structural and Semantic Perturbations19 Apr 2025 0 repositories listed
-
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers19 Apr 2025 0 repositories listed
-
Unreal Robotics Lab: A High-Fidelity Robotics Simulator with Advanced Physics and Rendering19 Apr 2025 0 repositories listed
-
Integrated Super-resolution Sensing and Symbiotic Communication with 3D Sparse MIMO for Low-Altitude UAV Swarm18 Apr 2025 0 repositories listed
-
OpenDeception: Benchmarking and Investigating AI Deceptive Behaviors via Open-ended Interaction Simulation18 Apr 2025 0 repositories listed
-
ALT: A Python Package for Lightweight Feature Representation in Time Series Classification17 Apr 2025 0 repositories listed
-
Benchmarking Multi-National Value Alignment for Large Language Models17 Apr 2025 0 repositories listed
-
Enhancing Explainability and Reliable Decision-Making in Particle Swarm Optimization through Communication Topologies17 Apr 2025 0 repositories listed
-
Featuremetric benchmarking: Quantum computer benchmarks based on circuit features17 Apr 2025 0 repositories listed
-
Local Data Quantity-Aware Weighted Averaging for Federated Learning with Dishonest Clients17 Apr 2025 0 repositories listed
-
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models17 Apr 2025 0 repositories listed
-
Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios16 Apr 2025 0 repositories listed
-
Benchmarking Mutual Information-based Loss Functions in Federated Learning16 Apr 2025 0 repositories listed
-
pix2pockets: Shot Suggestions in 8-Ball Pool from a Single Image in the Wild16 Apr 2025 0 repositories listed
-
Power Line Communication vs. Talkative Power Conversion: A Benchmarking Study16 Apr 2025 0 repositories listed
-
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions16 Apr 2025 0 repositories listed
-
BEACON: A Benchmark for Efficient and Accurate Counting of Subgraphs15 Apr 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.