Browse State-of-the-Art › Benchmarking › Papers, page 38
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 38 of 56: papers 3,701 to 3,800 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA22 Jul 2024 0 repositories listed
-
InLUT3D: Challenging real indoor dataset for point cloud analysis22 Jul 2024 0 repositories listed
-
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation22 Jul 2024 0 repositories listed
-
Unlocking the Potential: Benchmarking Large Language Models in Water Engineering and Research22 Jul 2024 0 repositories listed
-
Non-Reference Quality Assessment for Medical Imaging: Application to Synthetic Brain MRIs20 Jul 2024 0 repositories listed
-
Benchmarking deep learning models for bearing fault diagnosis using the CWRU dataset: A multi-label approach19 Jul 2024 0 repositories listed
-
OCTrack: Benchmarking the Open-Corpus Multi-Object Tracking19 Jul 2024 0 repositories listed
-
Realistic Evaluation of Test-Time Adaptation Algorithms: Unsupervised Hyperparameter Selection19 Jul 2024 0 repositories listed
-
SHS: Scorpion Hunting Strategy Swarm Algorithm19 Jul 2024 0 repositories listed
-
Vision-Based Power Line Cables and Pylons Detection for Low Flying Aircraft19 Jul 2024 0 repositories listed
-
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance18 Jul 2024 0 repositories listed
-
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle18 Jul 2024 0 repositories listed
-
RT-Pose: A 4D Radar Tensor-based 3D Human Pose Estimation and Localization Benchmark18 Jul 2024 0 repositories listed
-
Comprehensive Review and Empirical Evaluation of Causal Discovery Algorithms for Numerical Data17 Jul 2024 0 repositories listed
-
FETCH: A Memory-Efficient Replay Approach for Continual Learning in Image Classification17 Jul 2024 0 repositories listed
-
HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects17 Jul 2024 0 repositories listed
-
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?17 Jul 2024 0 repositories listed
-
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification16 Jul 2024 0 repositories listed
-
AstroMLab 1: Who Wins Astronomy Jeopardy!?15 Jul 2024 0 repositories listed
-
Benchmarking Vision Language Models for Cultural Understanding15 Jul 2024 0 repositories listed
-
ConvBench: A Comprehensive Benchmark for 2D Convolution Primitive Evaluation15 Jul 2024 0 repositories listed
-
On Machine Learning Approaches for Protein-Ligand Binding Affinity Prediction15 Jul 2024 0 repositories listed
-
Experimental Benchmarking of Energy-saving Sub-Optimal Sliding Mode Control14 Jul 2024 0 repositories listed
-
Automated detection of gibbon calls from passive acoustic monitoring data using convolutional neural networks in the "torch for R" ecosystem13 Jul 2024 0 repositories listed
-
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs13 Jul 2024 0 repositories listed
-
A Comprehensive Survey on Retrieval Methods in Recommender Systems11 Jul 2024 0 repositories listed
-
Evaluating Nuanced Bias in Large Language Model Free Response Answers11 Jul 2024 0 repositories listed
-
Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models10 Jul 2024 0 repositories listed
-
How Aligned are Different Alignment Metrics?10 Jul 2024 0 repositories listed
-
Analyzing the Effectiveness of Listwise Reranking with Positional Invariance on Temporal Generalizability9 Jul 2024 0 repositories listed
-
SPINEX-Clustering: Similarity-based Predictions with Explainable Neighbors Exploration for Clustering Problems9 Jul 2024 0 repositories listed
-
A Benchmark for Multi-speaker Anonymization8 Jul 2024 0 repositories listed
-
GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation8 Jul 2024 0 repositories listed
-
MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition8 Jul 2024 0 repositories listed
-
TARGO: Benchmarking Target-driven Object Grasping under Occlusions8 Jul 2024 0 repositories listed
-
Benchmarking GNNs Using Lightning Network Data5 Jul 2024 0 repositories listed
-
From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano5 Jul 2024 0 repositories listed
-
Towards Stable 3D Object Detection5 Jul 2024 0 repositories listed
-
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation4 Jul 2024 0 repositories listed
-
Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms3 Jul 2024 0 repositories listed
-
Open foundation models for Azerbaijani language2 Jul 2024 0 repositories listed
-
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations2 Jul 2024 0 repositories listed
-
EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting1 Jul 2024 0 repositories listed
-
MIRAI: Evaluating LLM Agents for Event Forecasting1 Jul 2024 0 repositories listed
-
Modified CMA-ES Algorithm for Multi-Modal Optimization: Incorporating Niching Strategies and Dynamic Adaptation Mechanism1 Jul 2024 0 repositories listed
-
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions1 Jul 2024 0 repositories listed
-
Task-oriented Over-the-air Computation for Edge-device Co-inference with Balanced Classification Accuracy1 Jul 2024 0 repositories listed
-
Commute Graph Neural Networks30 Jun 2024 0 repositories listed
-
GenderBias-VL: Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing30 Jun 2024 0 repositories listed
-
PerSEval: Assessing Personalization in Text Summarizers29 Jun 2024 0 repositories listed
-
Benchmarking M6 Competitors: An Analysis of Financial Metrics and Discussion of Incentives27 Jun 2024 0 repositories listed
-
Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges27 Jun 2024 0 repositories listed
-
Evaluating and Benchmarking Foundation Models for Earth Observation and Geospatial AI26 Jun 2024 0 repositories listed
-
Quantum-tunnelling deep neural network for optical illusion recognition26 Jun 2024 0 repositories listed
-
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis26 Jun 2024 0 repositories listed
-
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models25 Jun 2024 0 repositories listed
-
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making25 Jun 2024 0 repositories listed
-
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language25 Jun 2024 0 repositories listed
-
25 Jun 2024 0 repositories listed Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems25 Jun 2024 0 repositories listed
-
CATBench: A Compiler Autotuning Benchmarking Suite for Black-box Optimization24 Jun 2024 0 repositories listed
-
MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models24 Jun 2024 0 repositories listed
-
PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs24 Jun 2024 0 repositories listed
-
GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets23 Jun 2024 0 repositories listed
-
Position: Benchmarking is Limited in Reinforcement Learning Research23 Jun 2024 0 repositories listed
-
CaT-BENCH: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans22 Jun 2024 0 repositories listed
-
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents21 Jun 2024 0 repositories listed
-
Sports Intelligence: Assessing the Sports Understanding Capabilities of Language Models through Question Answering from Text to Video21 Jun 2024 0 repositories listed
-
Benchmarking Monocular 3D Dog Pose Estimation Using In-The-Wild Motion Capture Data20 Jun 2024 0 repositories listed
-
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines20 Jun 2024 0 repositories listed
-
DASB -- Discrete Audio and Speech Benchmark20 Jun 2024 0 repositories listed
-
Improving Expert Radiology Report Summarization by Prompting Large Language Models with a Layperson Summary20 Jun 2024 0 repositories listed
-
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions20 Jun 2024 0 repositories listed
-
Resource-efficient Medical Image Analysis with Self-adapting Forward-Forward Networks20 Jun 2024 0 repositories listed
-
Comparison of Open-Source and Proprietary LLMs for Machine Reading Comprehension: A Practical Analysis for Industrial Applications19 Jun 2024 0 repositories listed
-
Enhancing Distractor Generation for Multiple-Choice Questions with Retrieval Augmented Pretraining and Knowledge Graph Integration19 Jun 2024 0 repositories listed
-
Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models19 Jun 2024 0 repositories listed
-
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective19 Jun 2024 0 repositories listed
-
Exploring and Benchmarking the Planning Capabilities of Large Language Models18 Jun 2024 0 repositories listed
-
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance18 Jun 2024 0 repositories listed
-
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts18 Jun 2024 0 repositories listed
-
A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models17 Jun 2024 0 repositories listed
-
Benchmarking of LLM Detection: Comparing Two Competing Approaches17 Jun 2024 0 repositories listed
-
InternalInspector I²: Robust Confidence Estimation in LLMs through Internal States17 Jun 2024 0 repositories listed
-
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models17 Jun 2024 0 repositories listed
-
The Liouville Generator for Producing Integrable Expressions17 Jun 2024 0 repositories listed
-
Unleashing OpenTitan's Potential: a Silicon-Ready Embedded Secure Element for Root of Trust and Cryptographic Offloading17 Jun 2024 0 repositories listed
-
Benchmarking Out-of-Distribution Generalization Capabilities of DNN-based Encoding Models for the Ventral Visual Cortex16 Jun 2024 0 repositories listed
-
Evaluating the Performance of Large Language Models via Debates16 Jun 2024 0 repositories listed
-
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning16 Jun 2024 0 repositories listed
-
GANmut: Generating and Modifying Facial Expressions16 Jun 2024 0 repositories listed
-
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment16 Jun 2024 0 repositories listed
-
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences16 Jun 2024 0 repositories listed
-
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results15 Jun 2024 0 repositories listed
-
Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming14 Jun 2024 0 repositories listed
-
Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework14 Jun 2024 0 repositories listed
-
On the Evaluation of Speech Foundation Models for Spoken Language Understanding14 Jun 2024 0 repositories listed
-
A Review of 315 Benchmark and Test Functions for Machine Learning Optimization Algorithms and Metaheuristics with Mathematical and Visual Descriptions13 Jun 2024 0 repositories listed
-
Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition13 Jun 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.