Browse State-of-the-Art › Benchmarking › Papers, page 32
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 32 of 56: papers 3,101 to 3,200 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Advancing Human-Machine Teaming: Concepts, Challenges, and Applications16 Mar 2025 0 repositories listed
-
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era16 Mar 2025 0 repositories listed
-
Genicious: Contextual Few-shot Prompting for Insights Discovery15 Mar 2025 0 repositories listed
-
Dataset Properties Shape the Success of Neuroimaging-Based Patient Stratification: A Benchmarking Analysis Across Clustering Algorithms15 Mar 2025 0 repositories listed
-
Language Models for Automated Classification of Brain MRI Reports and Growth Chart Generation15 Mar 2025 0 repositories listed
-
Challenges and Advancements in Modeling Shock Fronts with Physics-Informed Neural Networks: A Review and Benchmarking Study14 Mar 2025 0 repositories listed
-
Dynamic Obstacle Avoidance with Bounded Rationality Adversarial Reinforcement Learning14 Mar 2025 0 repositories listed
-
Enhancing Hand Palm Motion Gesture Recognition by Eliminating Reference Frame Bias via Frame-Invariant Similarity Measures14 Mar 2025 0 repositories listed
-
Heterogeneous graph neural networks for species distribution modeling14 Mar 2025 0 repositories listed
-
InverseBench: Benchmarking Plug-and-Play Diffusion Priors for Inverse Problems in Physical Sciences14 Mar 2025 0 repositories listed
-
LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama14 Mar 2025 0 repositories listed
-
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation14 Mar 2025 0 repositories listed
-
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning14 Mar 2025 0 repositories listed
-
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity14 Mar 2025 0 repositories listed
-
DarkBench: Benchmarking Dark Patterns in Large Language Models13 Mar 2025 0 repositories listed
-
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content13 Mar 2025 0 repositories listed
-
TIME: Temporal-sensitive Multi-dimensional Instruction Tuning and Benchmarking for Video-LLMs13 Mar 2025 0 repositories listed
-
CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding12 Mar 2025 0 repositories listed
-
MarineGym: A High-Performance Reinforcement Learning Platform for Underwater Robotics12 Mar 2025 0 repositories listed
-
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models12 Mar 2025 0 repositories listed
-
Comprehensive Benchmarking of Machine Learning Methods for Risk Prediction Modelling from Large-Scale Survival Data: A UK Biobank Study11 Mar 2025 0 repositories listed
-
Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking11 Mar 2025 0 repositories listed
-
ResBench: Benchmarking LLM-Generated FPGA Designs with Resource Awareness11 Mar 2025 0 repositories listed
-
Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies10 Mar 2025 0 repositories listed
-
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models10 Mar 2025 0 repositories listed
-
Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems9 Mar 2025 0 repositories listed
-
General Scales Unlock AI Evaluation with Explanatory and Predictive Power9 Mar 2025 0 repositories listed
-
Is Your Benchmark (Still) Useful? Dynamic Benchmarking for Code Language Models9 Mar 2025 0 repositories listed
-
Steerable Pyramid Weighted Loss: Multi-Scale Adaptive Weighting for Semantic Segmentation9 Mar 2025 0 repositories listed
-
Removing Multiple Hybrid Adverse Weather in Video via a Unified Model8 Mar 2025 0 repositories listed
-
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces8 Mar 2025 0 repositories listed
-
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Enhancement Protocol7 Mar 2025 0 repositories listed
-
Benchmarking LLMs in Recommendation Tasks: A Comparative Evaluation with Conventional Recommenders7 Mar 2025 0 repositories listed
-
FinTMMBench: Benchmarking Temporal-Aware Multi-Modal RAG in Finance7 Mar 2025 0 repositories listed
-
Understanding the Limits of Lifelong Knowledge Editing in LLMs7 Mar 2025 0 repositories listed
-
Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms6 Mar 2025 0 repositories listed
-
Benchmarking Reasoning Robustness in Large Language Models6 Mar 2025 0 repositories listed
-
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination6 Mar 2025 0 repositories listed
-
Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering Datasets6 Mar 2025 0 repositories listed
-
Eventprop training for efficient neuromorphic applications6 Mar 2025 0 repositories listed
-
InfoSEM: A Deep Generative Model with Informative Priors for Gene Regulatory Network Inference6 Mar 2025 0 repositories listed
-
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges6 Mar 2025 0 repositories listed
-
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression6 Mar 2025 0 repositories listed
-
A2Perf: Real-World Autonomous Agents Benchmark4 Mar 2025 0 repositories listed
-
Evaluation of Architectural Synthesis Using Generative AI4 Mar 2025 0 repositories listed
-
Optimizing open-domain question answering with graph-based retrieval augmented generation4 Mar 2025 0 repositories listed
-
Technical report of a DMD-based Characterization Method for Vision Sensors4 Mar 2025 0 repositories listed
-
Multi-Agent Reinforcement Learning with Long-Term Performance Objectives for Service Workforce Optimization3 Mar 2025 0 repositories listed
-
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models3 Mar 2025 0 repositories listed
-
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics3 Mar 2025 0 repositories listed
-
Towards Efficient Educational Chatbots: Benchmarking RAG Frameworks2 Mar 2025 0 repositories listed
-
MAPS: Multi-Fidelity AI-Augmented Photonic Simulation and Inverse Design Infrastructure2 Mar 2025 0 repositories listed
-
FunBench: Benchmarking Fundus Reading Skills of MLLMs2 Mar 2025 0 repositories listed
-
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information1 Mar 2025 0 repositories listed
-
Large Language Model-Based Benchmarking Experiment Settings for Evolutionary Multi-Objective Optimization28 Feb 2025 0 repositories listed
-
ProBench: Benchmarking Large Language Models in Competitive Programming28 Feb 2025 0 repositories listed
-
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice28 Feb 2025 0 repositories listed
-
Solar Multimodal Transformer: Intraday Solar Irradiance Predictor using Public Cameras and Time Series28 Feb 2025 0 repositories listed
-
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments27 Feb 2025 0 repositories listed
-
MMSciBench: Benchmarking Language Models on Multimodal Scientific Problems27 Feb 2025 0 repositories listed
-
Agentic Mixture-of-Workflows for Multi-Modal Chemical Search26 Feb 2025 0 repositories listed
-
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv1026 Feb 2025 0 repositories listed
-
Is Your Paper Being Reviewed by an LLM? A New Benchmark Dataset and Approach for Detecting AI Text in Peer Review26 Feb 2025 0 repositories listed
-
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval26 Feb 2025 0 repositories listed
-
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors26 Feb 2025 0 repositories listed
-
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering26 Feb 2025 0 repositories listed
-
Modelling Regional Solar Photovoltaic Capacity in Great Britain26 Feb 2025 0 repositories listed
-
A Real-time Spatio-Temporal Trajectory Planner for Autonomous Vehicles with Semantic Graph Optimization25 Feb 2025 0 repositories listed
-
CayleyPy RL: Pathfinding and Reinforcement Learning on Cayley Graphs25 Feb 2025 0 repositories listed
-
OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation25 Feb 2025 0 repositories listed
-
Science Across Languages: Assessing LLM Multilingual Translation of Scientific Papers25 Feb 2025 0 repositories listed
-
Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement24 Feb 2025 0 repositories listed
-
Overconfident Oracles: Limitations of In Silico Sequence Design Benchmarking24 Feb 2025 0 repositories listed
-
SynthRAD2025 Grand Challenge dataset: generating synthetic CTs for radiotherapy24 Feb 2025 0 repositories listed
-
Benchmarking Online Object Trackers for Underwater Robot Position Locking Applications23 Feb 2025 0 repositories listed
-
On Neural Inertial Classification Networks for Pedestrian Activity Recognition23 Feb 2025 0 repositories listed
-
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs23 Feb 2025 0 repositories listed
-
VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models23 Feb 2025 0 repositories listed
-
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation21 Feb 2025 0 repositories listed
-
Methods and Trends in Detecting Generated Images: A Comprehensive Review21 Feb 2025 0 repositories listed
-
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models21 Feb 2025 0 repositories listed
-
Para-Lane: Multi-Lane Dataset Registering Parallel Scans for Benchmarking Novel View Synthesis21 Feb 2025 0 repositories listed
-
Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems20 Feb 2025 0 repositories listed
-
Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models20 Feb 2025 0 repositories listed
-
Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks20 Feb 2025 0 repositories listed
-
Probabilistic Robustness in Deep Learning: A Concise yet Comprehensive Guide20 Feb 2025 0 repositories listed
-
Reinforcement Learning with Graph Attention for Routing and Wavelength Assignment with Lightpath Reuse20 Feb 2025 0 repositories listed
-
Sentence Smith: Formally Controllable Text Transformation and its Application to Evaluation of Text Embedding Models20 Feb 2025 0 repositories listed
-
Statistical Scenario Modelling and Lookalike Distributions for Multi-Variate AI Risk20 Feb 2025 0 repositories listed
-
A Baseline Method for Removing Invisible Image Watermarks using Deep Image Prior19 Feb 2025 0 repositories listed
-
Benchmarking of Different YOLO Models for CAPTCHAs Detection and Classification19 Feb 2025 0 repositories listed
-
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking19 Feb 2025 0 repositories listed
-
Position: There are no Champions in Long-Term Time Series Forecasting19 Feb 2025 0 repositories listed
-
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare19 Feb 2025 0 repositories listed
-
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical Diagnostics18 Feb 2025 0 repositories listed
-
Benchmarking MedMNIST dataset on real quantum hardware18 Feb 2025 0 repositories listed
-
A new pathway to generative artificial intelligence by minimizing the maximum entropy18 Feb 2025 0 repositories listed
-
EquiBench: Benchmarking Large Language Models' Understanding of Program Semantics via Equivalence Checking18 Feb 2025 0 repositories listed
-
LLMPopcorn: An Empirical Study of LLMs as Assistants for Popular Micro-video Generation18 Feb 2025 0 repositories listed
-
Multilingual European Language Models: Benchmarking Approaches and Challenges18 Feb 2025 0 repositories listed