Browse State-of-the-Art › Benchmarking › Papers, page 31
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 31 of 56: papers 3,001 to 3,100 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items15 Apr 2025 0 repositories listed
-
Benchmarking Vision Language Models on German Factual Data15 Apr 2025 0 repositories listed
-
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives15 Apr 2025 0 repositories listed
-
E2E Parking Dataset: An Open Benchmark for End-to-End Autonomous Parking15 Apr 2025 0 repositories listed
-
GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDR15 Apr 2025 0 repositories listed
-
Benchmarking 3D Human Pose Estimation Models Under Occlusions14 Apr 2025 0 repositories listed
-
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design14 Apr 2025 0 repositories listed
-
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models14 Apr 2025 0 repositories listed
-
BoTTA: Benchmarking on-device Test Time Adaptation14 Apr 2025 0 repositories listed
-
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography14 Apr 2025 0 repositories listed
-
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts14 Apr 2025 0 repositories listed
-
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization14 Apr 2025 0 repositories listed
-
LMFormer: Lane based Motion Prediction Transformer14 Apr 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
SortBench: Benchmarking LLMs based on their ability to sort lists11 Apr 2025 0 repositories listed
-
TP-RAG: Benchmarking Retrieval-Augmented Large Language Model Agents for Spatiotemporal-Aware Travel Planning11 Apr 2025 0 repositories listed
-
Benchmarking Image Embeddings for E-Commerce: Evaluating Off-the Shelf Foundation Models, Fine-Tuning Strategies and Practical Trade-offs10 Apr 2025 0 repositories listed
-
Benchmarking Multi-Organ Segmentation Tools for Multi-Parametric T1-weighted Abdominal MRI10 Apr 2025 0 repositories listed
-
SydneyScapes: Image Segmentation for Australian Environments10 Apr 2025 0 repositories listed
-
A Roadmap for Improving Data Reliability and Sharing in Crosslinking Mass Spectrometry9 Apr 2025 0 repositories listed
-
Benchmarking Convolutional Neural Network and Graph Neural Network based Surrogate Models on a Real-World Car External Aerodynamics Dataset9 Apr 2025 0 repositories listed
-
Can Carbon-Aware Electric Load Shifting Reduce Emissions? An Equilibrium-Based Analysis9 Apr 2025 0 repositories listed
-
RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration9 Apr 2025 0 repositories listed
-
TabKAN: Advancing Tabular Data Analysis using Kolmogorov-Arnold Network9 Apr 2025 0 repositories listed
-
A Solid-State Nanopore Signal Generator for Training Machine Learning Models7 Apr 2025 0 repositories listed
-
Cross-functional transferability in universal machine learning interatomic potentials7 Apr 2025 0 repositories listed
-
Generative Adversarial Networks with Limited Data: A Survey and Benchmarking7 Apr 2025 0 repositories listed
-
Leveraging State Space Models in Long Range Genomics7 Apr 2025 0 repositories listed
-
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search7 Apr 2025 0 repositories listed
-
Riemannian Geometry for the classification of brain states with intracortical brain-computer interfaces7 Apr 2025 0 repositories listed
-
Towards Visual Text Grounding of Multimodal Large Language Model7 Apr 2025 0 repositories listed
-
Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams4 Apr 2025 0 repositories listed
-
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models4 Apr 2025 0 repositories listed
-
Point Cloud Objective Quality: Benchmarking Features and Quality Evaluation4 Apr 2025 0 repositories listed
-
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency4 Apr 2025 0 repositories listed
-
Towards a Unified Framework for Determining Conformational Ensembles of Disordered Proteins4 Apr 2025 0 repositories listed
-
Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-ray: Summary of the PENGWIN 2024 Challenge3 Apr 2025 0 repositories listed
-
Accelerating IoV Intrusion Detection: Benchmarking GPU-Accelerated vs CPU-Based ML Libraries2 Apr 2025 0 repositories listed
-
Benchmarking the Spatial Robustness of DNNs via Natural and Adversarial Localized Corruptions2 Apr 2025 0 repositories listed
-
Better Bill GPT: Comparing Large Language Models against Legal Invoice Reviewers2 Apr 2025 0 repositories listed
-
FIORD: A Fisheye Indoor-Outdoor Dataset with LIDAR Ground Truth for 3D Scene Reconstruction and Benchmarking2 Apr 2025 0 repositories listed
-
Global Rice Multi-Class Segmentation Dataset (RiceSEG): A Comprehensive and Diverse High-Resolution RGB-Annotated Images for the Development and Benchmarking of Rice Segmentation Algorithms2 Apr 2025 0 repositories listed
-
Horizon Scans can be accelerated using novel information retrieval and artificial intelligence tools2 Apr 2025 0 repositories listed
-
Proof of Humanity: A Multi-Layer Network Framework for Certifying Human-Originated Content in an AI-Dominated Internet2 Apr 2025 0 repositories listed
-
When Reasoning Meets Compression: Benchmarking Compressed Large Reasoning Models on Complex Reasoning Tasks2 Apr 2025 0 repositories listed
-
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models1 Apr 2025 0 repositories listed
-
Benchmarking Federated Machine Unlearning methods for Tabular Data1 Apr 2025 0 repositories listed
-
On Benchmarking Code LLMs for Android Malware Analysis1 Apr 2025 0 repositories listed
-
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models1 Apr 2025 0 repositories listed
-
Towards Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-critical Scenarios31 Mar 2025 0 repositories listed
-
Uni-Render: A Unified Accelerator for Real-Time Rendering Across Diverse Neural Renderers31 Mar 2025 0 repositories listed
-
Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models30 Mar 2025 0 repositories listed
-
Simple Feedfoward Neural Networks are Almost All You Need for Time Series Forecasting30 Mar 2025 0 repositories listed
-
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis29 Mar 2025 0 repositories listed
-
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation29 Mar 2025 0 repositories listed
-
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations29 Mar 2025 0 repositories listed
-
An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model28 Mar 2025 0 repositories listed
-
Assessing Foundation Models for Sea Ice Type Segmentation in Sentinel-1 SAR Imagery28 Mar 2025 0 repositories listed
-
Benchmarking Ultra-Low-Power μNPUs28 Mar 2025 0 repositories listed
-
Generalization Bias in Large Language Model Summarization of Scientific Research28 Mar 2025 0 repositories listed
-
LIM: Large Interpolator Model for Dynamic Reconstruction28 Mar 2025 0 repositories listed
-
SimBank: from Simulation to Solution in Prescriptive Process Monitoring28 Mar 2025 0 repositories listed
-
Benchmarking Deep Learning-Based Methods for Irradiance Nowcasting with Sky Images27 Mar 2025 0 repositories listed
-
Evaluating Text-to-Image Synthesis with a Conditional Fréchet Distance27 Mar 2025 0 repositories listed
-
GateLens: A Reasoning-Enhanced LLM Agent for Automotive Software Release Analytics27 Mar 2025 0 repositories listed
-
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition27 Mar 2025 0 repositories listed
-
Benchmarking Machine Learning Methods for Distributed Acoustic Sensing26 Mar 2025 0 repositories listed
-
CSPO: Cross-Market Synergistic Stock Price Movement Forecasting with Pseudo-volatility Optimization26 Mar 2025 0 repositories listed
-
RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy26 Mar 2025 0 repositories listed
-
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy25 Mar 2025 0 repositories listed
-
Reservoir Computing with a Single Oscillating Gas Bubble: Emphasizing the Chaotic Regime25 Mar 2025 0 repositories listed
-
Writing as a testbed for open ended agents25 Mar 2025 0 repositories listed
-
Benchmarking Burst Super-Resolution for Polarization Images: Noise Dataset and Analysis24 Mar 2025 0 repositories listed
-
Benchmarking Post-Hoc Unknown-Category Detection in Food Recognition24 Mar 2025 0 repositories listed
-
Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages24 Mar 2025 0 repositories listed
-
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation24 Mar 2025 0 repositories listed
-
A Study on Neuro-Symbolic Artificial Intelligence: Healthcare Perspectives23 Mar 2025 0 repositories listed
-
Regularization of ML models for Earth systems by using longer model timesteps23 Mar 2025 0 repositories listed
-
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering23 Mar 2025 0 repositories listed
-
Benchmark Dataset for Pore-Scale CO2-Water Interaction22 Mar 2025 0 repositories listed
-
CardioTabNet: A Novel Hybrid Transformer Model for Heart Disease Prediction using Tabular Medical Data22 Mar 2025 0 repositories listed
-
CausalRivers -- Scaling up benchmarking of causal discovery for real-world time-series21 Mar 2025 0 repositories listed
-
A Statistical Analysis for Per-Instance Evaluation of Stochastic Optimizers: How Many Repeats Are Enough?20 Mar 2025 0 repositories listed
-
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs20 Mar 2025 0 repositories listed
-
ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph20 Mar 2025 0 repositories listed
-
Empirical Analysis of Privacy-Fairness-Accuracy Trade-offs in Federated Learning: A Step Towards Responsible AI20 Mar 2025 0 repositories listed
-
Benchmarking Large Language Models for Handwritten Text Recognition19 Mar 2025 0 repositories listed
-
Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks19 Mar 2025 0 repositories listed
-
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding19 Mar 2025 0 repositories listed
-
ImputeGAP: A Comprehensive Library for Time Series Imputation19 Mar 2025 0 repositories listed
-
Kolmogorov-Arnold Network for Transistor Compact Modeling19 Mar 2025 0 repositories listed
-
SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban Meshes19 Mar 2025 0 repositories listed
-
ConSCompF: Consistency-focused Similarity Comparison Framework for Generative Large Language Models18 Mar 2025 0 repositories listed
-
COPA: Comparing the Incomparable to Explore the Pareto Front18 Mar 2025 0 repositories listed
-
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack18 Mar 2025 0 repositories listed
-
HA-VLN: A Benchmark for Human-Aware Navigation in Discrete-Continuous Environments with Dynamic Multi-Human Interactions, Real-World Validation, and an Open Leaderboard18 Mar 2025 0 repositories listed
-
Organ-aware Multi-scale Medical Image Segmentation Using Text Prompt Engineering18 Mar 2025 0 repositories listed
-
Stable Virtual Camera: Generative View Synthesis with Diffusion Models18 Mar 2025 0 repositories listed
-
VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination17 Mar 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.