Browse State-of-the-Art › Benchmarking › Papers, page 34
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 34 of 56: papers 3,301 to 3,400 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI13 Jan 2025 0 repositories listed
-
Benchmarking YOLOv8 for Optimal Crack Detection in Civil Infrastructure12 Jan 2025 0 repositories listed
-
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition10 Jan 2025 0 repositories listed
-
AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI9 Jan 2025 0 repositories listed
-
CallNavi, A Challenge and Empirical Study on LLM Function Calling and Routing9 Jan 2025 0 repositories listed
-
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning9 Jan 2025 0 repositories listed
-
Large Physics Models: Towards a collaborative approach with Large Language Models and Foundation Models9 Jan 2025 0 repositories listed
-
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation9 Jan 2025 0 repositories listed
-
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization8 Jan 2025 0 repositories listed
-
An Analysis of Model Robustness across Concurrent Distribution Shifts8 Jan 2025 0 repositories listed
-
Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks8 Jan 2025 0 repositories listed
-
Machine Learning for Identifying Grain Boundaries in Scanning Electron Microscopy (SEM) Images of Nanoparticle Superlattices7 Jan 2025 0 repositories listed
-
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding7 Jan 2025 0 repositories listed
-
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models6 Jan 2025 0 repositories listed
-
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input6 Jan 2025 0 repositories listed
-
AI-Powered Cow Detection in Complex Farm Environments3 Jan 2025 0 repositories listed
-
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation3 Jan 2025 0 repositories listed
-
PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents3 Jan 2025 0 repositories listed
-
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture3 Jan 2025 0 repositories listed
-
Benchmarking Constraint-Based Bayesian Structure Learning Algorithms: Role of Network Topology2 Jan 2025 0 repositories listed
-
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings2 Jan 2025 0 repositories listed
-
MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception2 Jan 2025 0 repositories listed
-
State-of-the-art AI-based Learning Approaches for Deepfake Generation and Detection, Analyzing Opportunities, Threading through Pros, Cons, and Future Prospects2 Jan 2025 0 repositories listed
-
TabTreeFormer: Tabular Data Generation Using Hybrid Tree-Transformer2 Jan 2025 0 repositories listed
-
CheXwhatsApp: A Dataset for Exploring Challenges in the Diagnosis of Chest X-rays through Mobile Devices1 Jan 2025 0 repositories listed
-
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools1 Jan 2025 0 repositories listed
-
CroCoDL: Cross-device Collaborative Dataset for Localization1 Jan 2025 0 repositories listed
-
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation1 Jan 2025 0 repositories listed
-
On the Utility of Equivariance and Symmetry Breaking in Deep Learning Architectures on Point Clouds1 Jan 2025 0 repositories listed
-
Segmenting Maxillofacial Structures in CBCT Volumes1 Jan 2025 0 repositories listed
-
Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion Models1 Jan 2025 0 repositories listed
-
Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedback1 Jan 2025 0 repositories listed
-
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation1 Jan 2025 0 repositories listed
-
A review of faithfulness metrics for hallucination assessment in Large Language Models31 Dec 2024 0 repositories listed
-
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects31 Dec 2024 0 repositories listed
-
Geometry Matters: Benchmarking Scientific ML Approaches for Flow Prediction around Complex Geometries31 Dec 2024 0 repositories listed
-
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing30 Dec 2024 0 repositories listed
-
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity30 Dec 2024 0 repositories listed
-
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI30 Dec 2024 0 repositories listed
-
Stratify: Unifying Multi-Step Forecasting Strategies29 Dec 2024 0 repositories listed
-
Towards Ideal Temporal Graph Neural Networks: Evaluations and Conclusions after 10,000 GPU Hours28 Dec 2024 0 repositories listed
-
Machine Generated Product Advertisements: Benchmarking LLMs Against Human Performance27 Dec 2024 0 repositories listed
-
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study25 Dec 2024 0 repositories listed
-
Re-assessing ImageNet: How aligned is its single-label assumption with its multi-label nature?24 Dec 2024 0 repositories listed
-
The Jungle of Generative Drug Discovery: Traps, Treasures, and Ways Out24 Dec 2024 0 repositories listed
-
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding23 Dec 2024 0 repositories listed
-
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations23 Dec 2024 0 repositories listed
-
Multimodal Deep Reinforcement Learning for Portfolio Optimization23 Dec 2024 0 repositories listed
-
SCBench: A Sports Commentary Benchmark for Video LLMs23 Dec 2024 0 repositories listed
-
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs23 Dec 2024 0 repositories listed
-
Patherea: Cell Detection and Classification for the 2020s21 Dec 2024 0 repositories listed
-
Benchmarking LLMs and SLMs for patient reported outcomes20 Dec 2024 0 repositories listed
-
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain20 Dec 2024 0 repositories listed
-
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage20 Dec 2024 0 repositories listed
-
Generation of Large District Heating System Models Using Open-Source Data and Tools: An Exemplary Workflow18 Dec 2024 0 repositories listed
-
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning18 Dec 2024 0 repositories listed
-
A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI17 Dec 2024 0 repositories listed
-
AI PERSONA: Towards Life-long Personalization of LLMs17 Dec 2024 0 repositories listed
-
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System17 Dec 2024 0 repositories listed
-
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration17 Dec 2024 0 repositories listed
-
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models17 Dec 2024 0 repositories listed
-
Selective Shot Learning for Code Explanation17 Dec 2024 0 repositories listed
-
ShiftedBronzes: Benchmarking and Analysis of Domain Fine-Grained Classification in Open-World Settings17 Dec 2024 0 repositories listed
-
How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games16 Dec 2024 0 repositories listed
-
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension16 Dec 2024 0 repositories listed
-
Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D Generation15 Dec 2024 0 repositories listed
-
Sequence-Level Leakage Risk of Training Data in Large Language Models15 Dec 2024 0 repositories listed
-
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries14 Dec 2024 0 repositories listed
-
Benchmarking large language models for materials synthesis: the case of atomic layer deposition13 Dec 2024 0 repositories listed
-
Benchmarking Table Comprehension In The Wild13 Dec 2024 0 repositories listed
-
CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems13 Dec 2024 0 repositories listed
-
Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction12 Dec 2024 0 repositories listed
-
Benchmarking of GPU-optimized Quantum-Inspired Evolutionary Optimization Algorithm using Functional Analysis12 Dec 2024 0 repositories listed
-
JuStRank: Benchmarking LLM Judges for System Ranking12 Dec 2024 0 repositories listed
-
Benchmarking learned algorithms for computed tomography image reconstruction tasks11 Dec 2024 0 repositories listed
-
Koopman Theory-Inspired Method for Learning Time Advancement Operators in Unstable Flame Front Evolution11 Dec 2024 0 repositories listed
-
LCFO: Long Context and Long Form Output Dataset and Benchmarking11 Dec 2024 0 repositories listed
-
Benchmarking Vision-Based Object Tracking for USVs in Complex Maritime Environments10 Dec 2024 0 repositories listed
-
Light Field Image Quality Assessment With Auxiliary Learning Based on Depthwise and Anglewise Separable Convolutions10 Dec 2024 0 repositories listed
-
MO-IOHinspector: Anytime Benchmarking of Multi-Objective Algorithms using IOHprofiler10 Dec 2024 0 repositories listed
-
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems10 Dec 2024 0 repositories listed
-
Towards Graph Foundation Models: A Study on the Generalization of Positional and Structural Encodings10 Dec 2024 0 repositories listed
-
Diff5T: Benchmarking Human Brain Diffusion MRI with an Extensive 5.0 Tesla K-Space and Spatial Dataset9 Dec 2024 0 repositories listed
-
How Certain are Uncertainty Estimates? Three Novel Earth Observation Datasets for Benchmarking Uncertainty Quantification in Machine Learning9 Dec 2024 0 repositories listed
-
Is Self-Supervision Enough? Benchmarking Foundation Models Against End-to-End Training for Mitotic Figure Classification9 Dec 2024 0 repositories listed
-
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions9 Dec 2024 0 repositories listed
-
On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events9 Dec 2024 0 repositories listed
-
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities9 Dec 2024 0 repositories listed
-
Vision-Based Deep Reinforcement Learning of UAV Autonomous Navigation Using Privileged Information9 Dec 2024 0 repositories listed
-
Evaluating Robustness of LLMs on Crisis-Related Microblogs across Events, Information Types, and Linguistic Features8 Dec 2024 0 repositories listed
-
Thermal Image-based Fault Diagnosis in Induction Machines via Self-Organized Operational Neural Networks8 Dec 2024 0 repositories listed
-
A Dataset Similarity Evaluation Framework for Wireless Communications and Sensing7 Dec 2024 0 repositories listed
-
ACT-Bench: Towards Action Controllable World Models for Autonomous Driving6 Dec 2024 0 repositories listed
-
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models6 Dec 2024 0 repositories listed
-
Learning Hidden Physics and System Parameters with Deep Operator Networks6 Dec 2024 0 repositories listed
-
MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects6 Dec 2024 0 repositories listed
-
MozzaVID: Mozzarella Volumetric Image Dataset6 Dec 2024 0 repositories listed
-
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage5 Dec 2024 0 repositories listed
-
Benchmarking and Enhancing Surgical Phase Recognition Models for Robotic-Assisted Esophagectomy5 Dec 2024 0 repositories listed
-
From Code to Play: Benchmarking Program Search for Games Using Large Language Models5 Dec 2024 0 repositories listed