Browse State-of-the-Art › Benchmarking › Papers, page 29
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 29 of 56: papers 2,801 to 2,900 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
MoE-Gyro: Self-Supervised Over-Range Reconstruction and Denoising for MEMS Gyroscopes27 May 2025 0 repositories listed
-
SOSBENCH: Benchmarking Safety Alignment on Scientific Knowledge27 May 2025 0 repositories listed
-
A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking26 May 2025 0 repositories listed
-
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems26 May 2025 0 repositories listed
-
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat26 May 2025 0 repositories listed
-
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages26 May 2025 0 repositories listed
-
EuroCon: Benchmarking Parliament Deliberation for Political Consensus Finding26 May 2025 0 repositories listed
-
PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology26 May 2025 0 repositories listed
-
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs26 May 2025 0 repositories listed
-
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs26 May 2025 0 repositories listed
-
Transformers in Protein: A Survey26 May 2025 0 repositories listed
-
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science25 May 2025 0 repositories listed
-
Benchmarking Large Language Models for Cyberbullying Detection in Real-World YouTube Comments25 May 2025 0 repositories listed
-
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research25 May 2025 0 repositories listed
-
EnvSDD: Benchmarking Environmental Sound Deepfake Detection25 May 2025 0 repositories listed
-
Retrieval-Augmented Generation for Service Discovery: Chunking Strategies and Benchmarking25 May 2025 0 repositories listed
-
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs25 May 2025 0 repositories listed
-
Where Paths Collide: A Comprehensive Survey of Classic and Learning-Based Multi-Agent Pathfinding25 May 2025 0 repositories listed
-
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation24 May 2025 0 repositories listed
-
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs24 May 2025 0 repositories listed
-
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation24 May 2025 0 repositories listed
-
LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Multi-Domain Reasoning Challenges24 May 2025 0 repositories listed
-
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models24 May 2025 0 repositories listed
-
So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection24 May 2025 0 repositories listed
-
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset24 May 2025 0 repositories listed
-
Benchmark for Antibody Binding Affinity Maturation and Design23 May 2025 0 repositories listed
-
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts23 May 2025 0 repositories listed
-
Is Single-View Mesh Reconstruction Ready for Robotics?23 May 2025 0 repositories listed
-
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation23 May 2025 0 repositories listed
-
PawPrint: Whose Footprints Are These? Identifying Animal Individuals by Their Footprints23 May 2025 0 repositories listed
-
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language23 May 2025 0 repositories listed
-
SEvoBench : A C++ Framework For Evolutionary Single-Objective Optimization Benchmarking23 May 2025 0 repositories listed
-
23 May 2025 0 repositories listed
-
BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Text22 May 2025 0 repositories listed
-
Benchmarking and Pushing the Multi-Bias Elimination Boundary of LLMs via Causal Effect Estimation-guided Debiasing22 May 2025 0 repositories listed
-
Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS222 May 2025 0 repositories listed
-
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research22 May 2025 0 repositories listed
-
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance22 May 2025 0 repositories listed
-
CUB: Benchmarking Context Utilisation Techniques for Language Models22 May 2025 0 repositories listed
-
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes22 May 2025 0 repositories listed
-
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs22 May 2025 0 repositories listed
-
Experimental robustness benchmark of quantum neural network on a superconducting quantum processor22 May 2025 0 repositories listed
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models22 May 2025 0 repositories listed
-
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models22 May 2025 0 repositories listed
-
MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries22 May 2025 0 repositories listed
-
MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks22 May 2025 0 repositories listed
-
When Safety Detectors Aren't Enough: A Stealthy and Effective Jailbreak Attack on LLMs via Steganographic Techniques22 May 2025 0 repositories listed
-
A Risk Taxonomy for Evaluating AI-Powered Psychotherapy Agents21 May 2025 0 repositories listed
-
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals21 May 2025 0 repositories listed
-
Benchmarking Chest X-ray Diagnosis Models Across Multinational Datasets21 May 2025 0 repositories listed
-
Benchmarking Energy and Latency in TinyML: A Novel Method for Resource-Constrained AI21 May 2025 0 repositories listed
-
Guidelines for the Quality Assessment of Energy-Aware NAS Benchmarks21 May 2025 0 repositories listed
-
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation21 May 2025 0 repositories listed
-
NEXT-EVAL: Next Evaluation of Traditional and LLM Web Data Record Extraction21 May 2025 0 repositories listed
-
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation21 May 2025 0 repositories listed
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems21 May 2025 0 repositories listed
-
Towards Zero-Shot Differential Morphing Attack Detection with Multimodal Large Language Models21 May 2025 0 repositories listed
-
UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning21 May 2025 0 repositories listed
-
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models21 May 2025 0 repositories listed
-
A Data-Driven Method to Identify IBRs with Dominant Participation in Sub-Synchronous Oscillations20 May 2025 0 repositories listed
-
Benchmarking data encoding methods in Quantum Machine Learning20 May 2025 0 repositories listed
-
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis20 May 2025 0 repositories listed
-
Explaining Unreliable Perception in Automated Driving: A Fuzzy-based Monitoring Approach20 May 2025 0 repositories listed
-
LLM-based Evaluation Policy Extraction for Ecological Modeling20 May 2025 0 repositories listed
-
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use20 May 2025 0 repositories listed
-
NavBench: A Unified Robotics Benchmark for Reinforcement Learning-Based Autonomous Navigation20 May 2025 0 repositories listed
-
NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI20 May 2025 0 repositories listed
-
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas20 May 2025 0 repositories listed
-
SlangDIT: Benchmarking LLMs in Interpretative Slang Translation20 May 2025 0 repositories listed
-
TransBench: Benchmarking Machine Translation for Industrial-Scale Applications20 May 2025 0 repositories listed
-
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations20 May 2025 0 repositories listed
-
A Comprehensive Benchmarking Platform for Deep Generative Models in Molecular Design19 May 2025 0 repositories listed
-
Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning19 May 2025 0 repositories listed
-
CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models19 May 2025 0 repositories listed
-
Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings19 May 2025 0 repositories listed
-
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference19 May 2025 0 repositories listed
-
LEXam: Benchmarking Legal Reasoning on 340 Law Exams19 May 2025 0 repositories listed
-
PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI19 May 2025 0 repositories listed
-
SzCORE as a benchmark: report from the seizure detection challenge at the 2025 AI in Epilepsy and Neurological Disorders Conference19 May 2025 0 repositories listed
-
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind18 May 2025 0 repositories listed
-
ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models18 May 2025 0 repositories listed
-
CompBench: Benchmarking Complex Instruction-guided Image Editing18 May 2025 0 repositories listed
-
Disambiguation in Conversational Question Answering in the Era of LLM: A Survey18 May 2025 0 repositories listed
-
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation17 May 2025 0 repositories listed
-
Machine Learning-Based Analysis of ECG and PCG Signals for Rheumatic Heart Disease Detection: A Scoping Review (2015-2025)17 May 2025 0 repositories listed
-
Benchmarking performance, explainability, and evaluation strategies of vision-language models for surgery: Challenges and opportunities16 May 2025 0 repositories listed
-
Relation Extraction Across Entire Books to Reconstruct Community Networks: The AffilKG Datasets16 May 2025 0 repositories listed
-
Visual Anomaly Detection under Complex View-Illumination Interplay: A Large-Scale Benchmark16 May 2025 0 repositories listed
-
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese16 May 2025 0 repositories listed
-
Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models16 May 2025 0 repositories listed
-
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems16 May 2025 0 repositories listed
-
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems16 May 2025 0 repositories listed
-
Benchmarking CFAR and CNN-based Peak Detection Algorithms in ISAC under Hardware Impairments16 May 2025 0 repositories listed
-
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale16 May 2025 0 repositories listed
-
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models16 May 2025 0 repositories listed
-
On the Evaluation of Engineering Artificial General Intelligence15 May 2025 0 repositories listed
-
Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization15 May 2025 0 repositories listed
-
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs15 May 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.