Browse State-of-the-Art › Benchmarking › Papers, page 35
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 35 of 56: papers 3,401 to 3,500 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models5 Dec 2024 0 repositories listed
-
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts5 Dec 2024 0 repositories listed
-
Uniform Discretized Integrated Gradients: An effective attribution based method for explaining large language models5 Dec 2024 0 repositories listed
-
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?4 Dec 2024 0 repositories listed
-
Benchmarking Attention Mechanisms and Consistency Regularization Semi-Supervised Learning for Post-Flood Building Damage Assessment in Satellite Images4 Dec 2024 0 repositories listed
-
Benchmarking Harmonized Tariff Schedule Classification Models4 Dec 2024 0 repositories listed
-
Benchmarking Pretrained Attention-based Models for Real-Time Recognition in Robot-Assisted Esophagectomy4 Dec 2024 0 repositories listed
-
Benchmarking terminology building capabilities of ChatGPT on an English-Russian Fashion Corpus4 Dec 2024 0 repositories listed
-
Benchmarking symbolic regression constant optimization schemes3 Dec 2024 0 repositories listed
-
OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations3 Dec 2024 0 repositories listed
-
Personalized Multimodal Large Language Models: A Survey3 Dec 2024 0 repositories listed
-
Single-Cell Omics Arena: A Benchmark Study for Large Language Models on Cell Type Annotation Using Single-Cell Data3 Dec 2024 0 repositories listed
-
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning3 Dec 2024 0 repositories listed
-
AI Benchmarks and Datasets for LLM Evaluation2 Dec 2024 0 repositories listed
-
Medchain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential Benchmarking2 Dec 2024 0 repositories listed
-
Understanding the World's Museums through Vision-Language Reasoning2 Dec 2024 0 repositories listed
-
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark29 Nov 2024 0 repositories listed
-
One-Shot Real-to-Sim via End-to-End Differentiable Simulation and Rendering29 Nov 2024 0 repositories listed
-
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks28 Nov 2024 0 repositories listed
-
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos28 Nov 2024 0 repositories listed
-
λ: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics28 Nov 2024 0 repositories listed
-
Benchmarking Agility and Reconfigurability in Satellite Systems for Tropical Cyclone Monitoring27 Nov 2024 0 repositories listed
-
Generating Diverse Synthetic Datasets for Evaluation of Real-life Recommender Systems27 Nov 2024 0 repositories listed
-
Agentic AI for Improving Precision in Identifying Contributions to Sustainable Development Goals26 Nov 2024 0 repositories listed
-
Evaluating Generative AI-Enhanced Content: A Conceptual Framework Using Qualitative, Quantitative, and Mixed-Methods Approaches26 Nov 2024 0 repositories listed
-
A Review of Bayesian Uncertainty Quantification in Deep Probabilistic Image Segmentation25 Nov 2024 0 repositories listed
-
Abnormality-Driven Representation Learning for Radiology Imaging25 Nov 2024 0 repositories listed
-
Performance Benchmarking of Psychomotor Skills Using Wearable Devices: An Application in Sport25 Nov 2024 0 repositories listed
-
Benchmarking Active Learning for NILM24 Nov 2024 0 repositories listed
-
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain23 Nov 2024 0 repositories listed
-
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains22 Nov 2024 0 repositories listed
-
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games20 Nov 2024 0 repositories listed
-
BelHouse3D: A Benchmark Dataset for Assessing Occlusion Robustness in 3D Point Cloud Semantic Segmentation20 Nov 2024 0 repositories listed
-
Benchmarking a wide range of optimisers for solving the Fermi-Hubbard model using the variational quantum eigensolver20 Nov 2024 0 repositories listed
-
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking20 Nov 2024 0 repositories listed
-
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis19 Nov 2024 0 repositories listed
-
The Moral Mind(s) of Large Language Models19 Nov 2024 0 repositories listed
-
Countering Backdoor Attacks in Image Recognition: A Survey and Evaluation of Mitigation Strategies17 Nov 2024 0 repositories listed
-
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML17 Nov 2024 0 repositories listed
-
FastDraft: How to Train Your Draft17 Nov 2024 0 repositories listed
-
Reinforcing Competitive Multi-Agents for Playing So Long Sucker17 Nov 2024 0 repositories listed
-
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level15 Nov 2024 0 repositories listed
-
Automated Coding of Communications in Collaborative Problem-solving Tasks Using ChatGPT15 Nov 2024 0 repositories listed
-
The Oxford Spires Dataset: Benchmarking Large-Scale LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods15 Nov 2024 0 repositories listed
-
The ParClusterers Benchmark Suite (PCBS): A Fine-Grained Analysis of Scalable Graph Clustering15 Nov 2024 0 repositories listed
-
WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking14 Nov 2024 0 repositories listed
-
A Survey on Vision Autoregressive Model13 Nov 2024 0 repositories listed
-
HyperFace: Generating Synthetic Face Recognition Datasets by Exploring Face Embedding Hypersphere13 Nov 2024 0 repositories listed
-
Evaluating the Generation of Spatial Relations in Text and Image Generative Models12 Nov 2024 0 repositories listed
-
11 Nov 2024 0 repositories listed
-
MolMiner: Towards Controllable, 3D-Aware, Fragment-Based Molecular Design10 Nov 2024 0 repositories listed
-
Low Dynamic Range for RIS-aided Bistatic Integrated Sensing and Communication9 Nov 2024 0 repositories listed
-
A Retrospective on the Robot Air Hockey Challenge: Benchmarking Robust, Reliable, and Safe Learning Techniques for Real-world Robotics8 Nov 2024 0 repositories listed
-
Benchmarking 3D multi-coil NC-PDNet MRI reconstruction8 Nov 2024 0 repositories listed
-
FactLens: Benchmarking Fine-Grained Fact Verification8 Nov 2024 0 repositories listed
-
Open-set object detection: towards unified problem formulation and benchmarking8 Nov 2024 0 repositories listed
-
Benchmarking Large Language Models with Integer Sequence Generation Tasks7 Nov 2024 0 repositories listed
-
Deep Learning Models for UAV-Assisted Bridge Inspection: A YOLO Benchmark Analysis7 Nov 2024 0 repositories listed
-
Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries7 Nov 2024 0 repositories listed
-
HandCraft: Anatomically Correct Restoration of Malformed Hands in Diffusion Generated Images7 Nov 2024 0 repositories listed
-
Learn to Solve Vehicle Routing Problems ASAP: A Neural Optimization Approach for Time-Constrained Vehicle Routing Problems with Finite Vehicle Fleet7 Nov 2024 0 repositories listed
-
Performance-Guided LLM Knowledge Distillation for Efficient Text Classification at Scale7 Nov 2024 0 repositories listed
-
Perspective on recent developments and challenges in regulatory and systems genomics7 Nov 2024 0 repositories listed
-
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding7 Nov 2024 0 repositories listed
-
Generating Synthetic Electronic Health Record (EHR) Data: A Review with Benchmarking6 Nov 2024 0 repositories listed
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level5 Nov 2024 0 repositories listed
-
SPINEX_ Symbolic Regression: Similarity-based Symbolic Regression with Explainable Neighbors Exploration5 Nov 2024 0 repositories listed
-
Benchmarking XAI Explanations with Human-Aligned Evaluations4 Nov 2024 0 repositories listed
-
Imagining and building wise machines: The centrality of AI metacognition4 Nov 2024 0 repositories listed
-
SinaTools: Open Source Toolkit for Arabic Natural Language Processing3 Nov 2024 0 repositories listed
-
Artificial Intelligence for Microbiology and Microbiome Research2 Nov 2024 0 repositories listed
-
Varco Arena: A Tournament Approach to Reference-Free Benchmarking Large Language Models2 Nov 2024 0 repositories listed
-
A Review of Reinforcement Learning in Financial Applications1 Nov 2024 0 repositories listed
-
Benchmarking Bias in Large Language Models during Role-Playing1 Nov 2024 0 repositories listed
-
Improving Few-Shot Cross-Domain Named Entity Recognition by Instruction Tuning a Word-Embedding based Retrieval Augmented Large Language Model1 Nov 2024 0 repositories listed
-
Modern, Efficient, and Differentiable Transport Equation Models using JAX: Applications to Population Balance Equations1 Nov 2024 0 repositories listed
-
Benchmark Data Repositories for Better Benchmarking31 Oct 2024 0 repositories listed
-
DexGraspNet 2.0: Learning Generative Dexterous Grasping in Large-scale Synthetic Cluttered Scenes30 Oct 2024 0 repositories listed
-
Evaluating Cultural and Social Awareness of LLM Web Agents30 Oct 2024 0 repositories listed
-
Low-Density 3D Point Cloud Classification30 Oct 2024 0 repositories listed
-
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning30 Oct 2024 0 repositories listed
-
Benchmarking LLM Guardrails in Handling Multilingual Toxicity29 Oct 2024 0 repositories listed
-
AI Cyber Risk Benchmark: Automated Exploitation Capabilities29 Oct 2024 0 repositories listed
-
SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset29 Oct 2024 0 repositories listed
-
BongLLaMA: LLaMA for Bangla Language28 Oct 2024 0 repositories listed
-
Exploring Capabilities of Time Series Foundation Models in Building Analytics28 Oct 2024 0 repositories listed
-
Hierarchical Knowledge Graph Construction from Images for Scalable E-Commerce28 Oct 2024 0 repositories listed
-
LLM-initialized Differentiable Causal Discovery28 Oct 2024 0 repositories listed
-
Project MPG: towards a generalized performance benchmark for LLM capabilities28 Oct 2024 0 repositories listed
-
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training28 Oct 2024 0 repositories listed
-
Multi-input Multi-output Loewner Framework for Vibration-based Damage Detection on a Trainer Jet26 Oct 2024 0 repositories listed
-
SFTrack: A Robust Scale and Motion Adaptive Algorithm for Tracking Small and Fast Moving Objects26 Oct 2024 0 repositories listed
-
A Survey of Small Language Models25 Oct 2024 0 repositories listed
-
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs25 Oct 2024 0 repositories listed
-
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding25 Oct 2024 0 repositories listed
-
OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery25 Oct 2024 0 repositories listed
-
Benchmarking Graph Learning for Drug-Drug Interaction Prediction24 Oct 2024 0 repositories listed
-
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems24 Oct 2024 0 repositories listed
-
Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling23 Oct 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.