Browse State-of-the-Art › Benchmarking › Papers, page 33
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 33 of 56: papers 3,201 to 3,300 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models18 Feb 2025 0 repositories listed
-
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation18 Feb 2025 0 repositories listed
-
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models17 Feb 2025 0 repositories listed
-
Ansatz-free Hamiltonian learning with Heisenberg-limited scaling17 Feb 2025 0 repositories listed
-
Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics17 Feb 2025 0 repositories listed
-
Energy-Conscious LLM Decoding: Impact of Text Generation Strategies on GPU Energy Consumption17 Feb 2025 0 repositories listed
-
Knowledge-aware contrastive heterogeneous molecular graph learning17 Feb 2025 0 repositories listed
-
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance17 Feb 2025 0 repositories listed
-
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment17 Feb 2025 0 repositories listed
-
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs16 Feb 2025 0 repositories listed
-
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking16 Feb 2025 0 repositories listed
-
User Profile with Large Language Models: Construction, Updating, and Benchmarking15 Feb 2025 0 repositories listed
-
Yesil o1 Pro: Evidence-Based AI Model for Health and Benchmarking in Clinical Decision Support15 Feb 2025 0 repositories listed
-
Benchmarking the rationality of AI decision making using the transitivity axiom14 Feb 2025 0 repositories listed
-
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow14 Feb 2025 0 repositories listed
-
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing14 Feb 2025 0 repositories listed
-
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?14 Feb 2025 0 repositories listed
-
A Survey on LLM-based News Recommender Systems13 Feb 2025 0 repositories listed
-
AT-Drone: Benchmarking Adaptive Teaming in Multi-Drone Pursuit13 Feb 2025 0 repositories listed
-
Beyond the Singular: The Essential Role of Multiple Generations in Effective Benchmark Evaluation and Analysis13 Feb 2025 0 repositories listed
-
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents13 Feb 2025 0 repositories listed
-
Machine learning for modelling unstructured grid data in computational physics: a review13 Feb 2025 0 repositories listed
-
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency13 Feb 2025 0 repositories listed
-
SkyRover: A Modular Simulator for Cross-Domain Pathfinding13 Feb 2025 0 repositories listed
-
Standardisation of Convex Ultrasound Data Through Geometric Analysis and Augmentation13 Feb 2025 0 repositories listed
-
Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors12 Feb 2025 0 repositories listed
-
Handwritten Text Recognition: A Survey12 Feb 2025 0 repositories listed
-
One-Shot Federated Learning with Classifier-Free Diffusion Models12 Feb 2025 0 repositories listed
-
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation10 Feb 2025 0 repositories listed
-
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories10 Feb 2025 0 repositories listed
-
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations10 Feb 2025 0 repositories listed
-
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models9 Feb 2025 0 repositories listed
-
Decoding Complexity: Intelligent Pattern Exploration with CHPDA (Context Aware Hybrid Pattern Detection Algorithm)9 Feb 2025 0 repositories listed
-
Surprise Potential as a Measure of Interactivity in Driving Scenarios8 Feb 2025 0 repositories listed
-
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization6 Feb 2025 0 repositories listed
-
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models6 Feb 2025 0 repositories listed
-
LUND-PROBE -- LUND Prostate Radiotherapy Open Benchmarking and Evaluation dataset6 Feb 2025 0 repositories listed
-
Verifiable Format Control for Large Language Model Generations6 Feb 2025 0 repositories listed
-
Benchmarking Time Series Forecasting Models: From Statistical Techniques to Foundation Models in Real-World Applications5 Feb 2025 0 repositories listed
-
Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials5 Feb 2025 0 repositories listed
-
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf5 Feb 2025 0 repositories listed
-
Optimal PMU Placement for Kalman Filtering of DAE Power System Models5 Feb 2025 0 repositories listed
-
xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods5 Feb 2025 0 repositories listed
-
Dynamic benchmarking framework for LLM-based conversational data capture4 Feb 2025 0 repositories listed
-
Evalita-LLM: Benchmarking Large Language Models on Italian4 Feb 2025 0 repositories listed
-
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models4 Feb 2025 0 repositories listed
-
LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation4 Feb 2025 0 repositories listed
-
EdgeMark: An Automation and Benchmarking System for Embedded Artificial Intelligence Tools3 Feb 2025 0 repositories listed
-
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation3 Feb 2025 0 repositories listed
-
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities3 Feb 2025 0 repositories listed
-
SE Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering3 Feb 2025 0 repositories listed
-
True Online TD-Replan(lambda) Achieving Planning through Replaying31 Jan 2025 0 repositories listed
-
Evolving Hard Maximum Cut Instances for Quantum Approximate Optimization Algorithms30 Jan 2025 0 repositories listed
-
Fine-tuning LLaMA 2 interference: a comparative study of language implementations for optimal efficiency30 Jan 2025 0 repositories listed
-
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding30 Jan 2025 0 repositories listed
-
Solving Urban Network Security Games: Learning Platform, Benchmark, and Challenge for AI Research29 Jan 2025 0 repositories listed
-
Benchmarking Quantum Convolutional Neural Networks for Signal Classification in Simulated Gamma-Ray Burst Detection28 Jan 2025 0 repositories listed
-
A Benchmarking Environment for Worker Flexibility in Flexible Job Shop Scheduling Problems27 Jan 2025 0 repositories listed
-
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding27 Jan 2025 0 repositories listed
-
Making Sense of Data in the Wild: Data Analysis Automation at Scale27 Jan 2025 0 repositories listed
-
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding27 Jan 2025 0 repositories listed
-
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation27 Jan 2025 0 repositories listed
-
Transfer of Knowledge through Reverse Annealing: A Preliminary Analysis of the Benefits and What to Share27 Jan 2025 0 repositories listed
-
Beyond Benchmarks: On The False Promise of AI Regulation26 Jan 2025 0 repositories listed
-
26 Jan 2025 0 repositories listed
-
Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?26 Jan 2025 0 repositories listed
-
Prompting ChatGPT for Chinese Learning as L2: A CEFR and EBCL Level Study25 Jan 2025 0 repositories listed
-
Benchmarking global optimization techniques for unmanned aerial vehicle path planning24 Jan 2025 0 repositories listed
-
Feature-based Evolutionary Diversity Optimization of Discriminating Instances for Chance-constrained Optimization Problems24 Jan 2025 0 repositories listed
-
The Karp Dataset24 Jan 2025 0 repositories listed
-
AEON: Adaptive Estimation of Instance-Dependent In-Distribution and Out-of-Distribution Label Noise for Robust Learning23 Jan 2025 0 repositories listed
-
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale23 Jan 2025 0 repositories listed
-
You Only Crash Once v2: Perceptually Consistent Strong Features for One-Stage Domain Adaptive Detection of Space Terrain23 Jan 2025 0 repositories listed
-
CHaRNet: Conditioned Heatmap Regression for Robust Dental Landmark Localization22 Jan 2025 0 repositories listed
-
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities22 Jan 2025 0 repositories listed
-
Leveraging LLMs to Create a Haptic Devices' Recommendation System22 Jan 2025 0 repositories listed
-
RAG-Reward: Optimizing RAG with Reward Modeling and RLHF22 Jan 2025 0 repositories listed
-
Benchmarking Generative AI for Scoring Medical Student Interviews in Objective Structured Clinical Examinations (OSCEs)21 Jan 2025 0 repositories listed
-
Benchmarking Randomized Optimization Algorithms on Binary, Permutation, and Combinatorial Problem Landscapes21 Jan 2025 0 repositories listed
-
Optimally-Weighted Maximum Mean Discrepancy Framework for Continual Learning21 Jan 2025 0 repositories listed
-
Algorithm Selection with Probing Trajectories: Benchmarking the Choice of Classifier Model20 Jan 2025 0 repositories listed
-
Benchmarking Large Language Models via Random Variables20 Jan 2025 0 repositories listed
-
Beyond the Hype: Benchmarking LLM-Evolved Heuristics for Bin Packing20 Jan 2025 0 repositories listed
-
An Interpretable Measure for Quantifying Predictive Dependence between Continuous Random Variables -- Extended Version18 Jan 2025 0 repositories listed
-
FORLAPS: An Innovative Data-Driven Reinforcement Learning Approach for Prescriptive Process Monitoring17 Jan 2025 0 repositories listed
-
Village-Net Clustering: A Rapid approach to Non-linear Unsupervised Clustering of High-Dimensional Data16 Jan 2025 0 repositories listed
-
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval15 Jan 2025 0 repositories listed
-
Cancer-Net PCa-Seg: Benchmarking Deep Learning Models for Prostate Cancer Segmentation Using Synthetic Correlated Diffusion Imaging15 Jan 2025 0 repositories listed
-
MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents15 Jan 2025 0 repositories listed
-
Off-policy Evaluation for Payments at Adyen15 Jan 2025 0 repositories listed
-
Similarity-Quantized Relative Difference Learning for Improved Molecular Activity Prediction15 Jan 2025 0 repositories listed
-
Benchmarking Classical, Deep, and Generative Models for Human Activity Recognition14 Jan 2025 0 repositories listed
-
Benchmarking Multimodal Models for Fine-Grained Image Analysis: A Comparative Study Across Diverse Visual Features14 Jan 2025 0 repositories listed
-
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving14 Jan 2025 0 repositories listed
-
Data-driven inventory management for new products: An adjusted Dyna-Q approach with transfer learning14 Jan 2025 0 repositories listed
-
Investigating Energy Efficiency and Performance Trade-offs in LLM Inference Across Tasks and DVFS Settings14 Jan 2025 0 repositories listed
-
Keras Sig: Efficient Path Signature Computation on GPU in Keras 314 Jan 2025 0 repositories listed
-
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles13 Jan 2025 0 repositories listed
-
Lessons From Red Teaming 100 Generative AI Products13 Jan 2025 0 repositories listed
-
The Paradox of Success in Evolutionary and Bioinspired Optimization: Revisiting Critical Issues, Key Studies, and Methodological Pathways13 Jan 2025 0 repositories listed