Browse State-of-the-Art › Benchmarking › Papers, page 28
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 28 of 56: papers 2,701 to 2,800 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Finance Language Model Evaluation (FLaME)18 Jun 2025 0 repositories listed
-
A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning17 Jun 2025 0 repositories listed
-
Egocentric Human-Object Interaction Detection: A New Benchmark and Method17 Jun 2025 0 repositories listed
-
PGLib-CO2: A Power Grid Library for Computing and Optimizing Carbon Emissions17 Jun 2025 0 repositories listed
-
Q2SAR: A Quantum Multiple Kernel Learning Approach for Drug Discovery17 Jun 2025 0 repositories listed
-
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects16 Jun 2025 0 repositories listed
-
Deep Diffusion Models and Unsupervised Hyperspectral Unmixing for Realistic Abundance Map Synthesis16 Jun 2025 0 repositories listed
-
Few-Shot Learning for Industrial Time Series: A Comparative Analysis Using the Example of Screw-Fastening Process Monitoring16 Jun 2025 0 repositories listed
-
JENGA: Object selection and pose estimation for robotic grasping from a stack16 Jun 2025 0 repositories listed
-
Robustness of Reinforcement Learning-Based Traffic Signal Control under Incidents: A Comparative Study16 Jun 2025 0 repositories listed
-
A large-scale, physically-based synthetic dataset for satellite pose estimation15 Jun 2025 0 repositories listed
-
Learning Best Paths in Quantum Networks14 Jun 2025 0 repositories listed
-
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables13 Jun 2025 0 repositories listed
-
crossMoDA Challenge: Evolution of Cross-Modality Domain Adaptation Techniques for Vestibular Schwannoma and Cochlea Segmentation from 2021 to 202313 Jun 2025 0 repositories listed
-
EconGym: A Scalable AI Testbed with Diverse Economic Tasks13 Jun 2025 0 repositories listed
-
SemanticST: Spatially Informed Semantic Graph Learning for Clustering, Integration, and Scalable Analysis of Spatial Transcriptomics13 Jun 2025 0 repositories listed
-
Temporal cross-validation impacts multivariate time series subsequence anomaly detection evaluation13 Jun 2025 0 repositories listed
-
Primender Sequence: A Novel Mathematical Construct for Testing Symbolic Inference and AI Reasoning12 Jun 2025 0 repositories listed
-
HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation12 Jun 2025 0 repositories listed
-
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics12 Jun 2025 0 repositories listed
-
Sum Rate Maximization for Pinching Antennas Assisted RSMA System With Multiple Waveguides12 Jun 2025 0 repositories listed
-
GRAIL: A Benchmark for GRaph ActIve Learning in Dynamic Sensing Environments11 Jun 2025 0 repositories listed
-
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents11 Jun 2025 0 repositories listed
-
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models11 Jun 2025 0 repositories listed
-
ICE-ID: A Novel Historical Census Data Benchmark Comparing NARS against LLMs, & a ML Ensemble on Longitudinal Identity Resolution11 Jun 2025 0 repositories listed
-
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models11 Jun 2025 0 repositories listed
-
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs11 Jun 2025 0 repositories listed
-
Solving excited states for long-range interacting trapped ions with neural networks10 Jun 2025 0 repositories listed
-
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP10 Jun 2025 0 repositories listed
-
Graph Attention-based Decentralized Actor-Critic for Dual-Objective Control of Multi-UAV Swarms10 Jun 2025 0 repositories listed
-
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens10 Jun 2025 0 repositories listed
-
Benchmarking Pre-Trained Time Series Models for Electricity Price Forecasting9 Jun 2025 0 repositories listed
-
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors9 Jun 2025 0 repositories listed
-
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework9 Jun 2025 0 repositories listed
-
Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech9 Jun 2025 0 repositories listed
-
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments9 Jun 2025 0 repositories listed
-
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding9 Jun 2025 0 repositories listed
-
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra9 Jun 2025 0 repositories listed
-
REMoH: A Reflective Evolution of Multi-objective Heuristics approach via Large Language Models9 Jun 2025 0 repositories listed
-
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents9 Jun 2025 0 repositories listed
-
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis9 Jun 2025 0 repositories listed
-
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures6 Jun 2025 0 repositories listed
-
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection6 Jun 2025 0 repositories listed
-
Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions6 Jun 2025 0 repositories listed
-
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques6 Jun 2025 0 repositories listed
-
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition5 Jun 2025 0 repositories listed
-
A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values5 Jun 2025 0 repositories listed
-
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs5 Jun 2025 0 repositories listed
-
Benchmarking Large Language Models on Homework Assessment in Circuit Analysis5 Jun 2025 0 repositories listed
-
CzechLynx: A Dataset for Individual Identification and Pose Estimation of the Eurasian Lynx5 Jun 2025 0 repositories listed
-
Design of intelligent proofreading system for English translation based on CNN and BERT5 Jun 2025 0 repositories listed
-
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models5 Jun 2025 0 repositories listed
-
FRED: The Florence RGB-Event Drone Dataset5 Jun 2025 0 repositories listed
-
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems5 Jun 2025 0 repositories listed
-
HoliSafe: Holistic Safety Benchmarking and Modeling with Safety Meta Token for Vision-Language Model5 Jun 2025 0 repositories listed
-
Refer to Anything with Vision-Language Prompts5 Jun 2025 0 repositories listed
-
Urania: Differentially Private Insights into AI Use5 Jun 2025 0 repositories listed
-
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos5 Jun 2025 0 repositories listed
-
Knowledge-guided Contextual Gene Set Analysis Using Large Language Models4 Jun 2025 0 repositories listed
-
CETBench: A Novel Dataset constructed via Transformations over Programs for Benchmarking LLMs for Code-Equivalence Checking4 Jun 2025 0 repositories listed
-
Curse of Slicing: Why Sliced Mutual Information is a Deceptive Measure of Statistical Dependence4 Jun 2025 0 repositories listed
-
Generating Automotive Code: Large Language Models for Software Development and Verification in Safety-Critical Systems4 Jun 2025 0 repositories listed
-
MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale4 Jun 2025 0 repositories listed
-
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP4 Jun 2025 0 repositories listed
-
Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset4 Jun 2025 0 repositories listed
-
AMLgentex: Mobilizing Data-Driven Research to Combat Money Laundering3 Jun 2025 0 repositories listed
-
FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models3 Jun 2025 0 repositories listed
-
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation3 Jun 2025 0 repositories listed
-
Tactile MNIST: Benchmarking Active Tactile Perception3 Jun 2025 0 repositories listed
-
Benchmarking Neural Speech Codec Intelligibility with SITool2 Jun 2025 0 repositories listed
-
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists2 Jun 2025 0 repositories listed
-
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents2 Jun 2025 0 repositories listed
-
Greening AI-enabled Systems with Software Engineering: A Research Agenda for Environmentally Sustainable AI Practices2 Jun 2025 0 repositories listed
-
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code2 Jun 2025 0 repositories listed
-
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?2 Jun 2025 0 repositories listed
-
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models1 Jun 2025 0 repositories listed
-
The iNaturalist Sounds Dataset31 May 2025 0 repositories listed
-
Automated Structured Radiology Report Generation30 May 2025 0 repositories listed
-
Benchmarking Foundation Models for Zero-Shot Biometric Tasks30 May 2025 0 repositories listed
-
Benchmarking Large Language Models for Cryptanalysis and Mismatched-Generalization30 May 2025 0 repositories listed
-
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents30 May 2025 0 repositories listed
-
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation30 May 2025 0 repositories listed
-
GenSpace: Benchmarking Spatially-Aware Image Generation30 May 2025 0 repositories listed
-
Geospatial Foundation Models to Enable Progress on Sustainable Development Goals30 May 2025 0 repositories listed
-
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models30 May 2025 0 repositories listed
-
Progressive Class-level Distillation30 May 2025 0 repositories listed
-
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking29 May 2025 0 repositories listed
-
Joint Phase Shift Optimization and Precoder Selection for RIS-Assisted 5G NR MIMO Systems29 May 2025 0 repositories listed
-
MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge29 May 2025 0 repositories listed
-
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation29 May 2025 0 repositories listed
-
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns29 May 2025 0 repositories listed
-
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates28 May 2025 0 repositories listed
-
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate28 May 2025 0 repositories listed
-
HelixDesign-Binder: A Scalable Production-Grade Platform for Binder Design Built on HelixFold328 May 2025 0 repositories listed
-
Jailbreak Distillation: Renewable Safety Benchmarking28 May 2025 0 repositories listed
-
PGLearn -- An Open-Source Learning Toolkit for Optimal Power Flow28 May 2025 0 repositories listed
-
TabularQGAN: A Quantum Generative Model for Tabular Data28 May 2025 0 repositories listed
-
Yambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval28 May 2025 0 repositories listed
-
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding27 May 2025 0 repositories listed
-
Gauss-Ramanujan Functions: Constructions, Properties, and Applications in Communications and Signal Processing27 May 2025 0 repositories listed