Browse State-of-the-Art › Benchmarking › Papers, page 37
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 37 of 56: papers 3,601 to 3,700 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG12 Sep 2024 0 repositories listed
-
Introducing CausalBench: A Flexible Benchmark Framework for Causal Analysis and Machine Learning12 Sep 2024 0 repositories listed
-
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots12 Sep 2024 0 repositories listed
-
The CLC-UKET Dataset: Benchmarking Case Outcome Prediction for the UK Employment Tribunal12 Sep 2024 0 repositories listed
-
The JPEG Pleno Learning-based Point Cloud Coding Standard: Serving Man and Machine12 Sep 2024 0 repositories listed
-
Benchmarking 2D Egocentric Hand Pose Datasets11 Sep 2024 0 repositories listed
-
Benchmarking and Validation of Sub-mW 30GHz VG-LNAs in 22nm FDSOI CMOS for 5G/6G Phased-Array Receivers11 Sep 2024 0 repositories listed
-
Understanding Foundation Models: Are We Back in 1924?11 Sep 2024 0 repositories listed
-
Benchmarking Sub-Genre Classification For Mainstage Dance Music10 Sep 2024 0 repositories listed
-
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding10 Sep 2024 0 repositories listed
-
Ransomware Detection Using Machine Learning in the Linux Kernel10 Sep 2024 0 repositories listed
-
VoiceWukong: Benchmarking Deepfake Voice Detection10 Sep 2024 0 repositories listed
-
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection9 Sep 2024 0 repositories listed
-
NeIn: Telling What You Don't Want9 Sep 2024 0 repositories listed
-
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E59 Sep 2024 0 repositories listed
-
RBoard: A Unified Platform for Reproducible and Reusable Recommender System Benchmarks9 Sep 2024 0 repositories listed
-
Selecting Differential Splicing Methods: Practical Considerations9 Sep 2024 0 repositories listed
-
Absolute Ranking: An Essential Normalization for Benchmarking Optimization Algorithms6 Sep 2024 0 repositories listed
-
Benchmarking Estimators for Natural Experiments: A Novel Dataset and a Doubly Robust Algorithm6 Sep 2024 0 repositories listed
-
Quantum Kernel Methods under Scrutiny: A Benchmarking Study6 Sep 2024 0 repositories listed
-
InfraLib: Enabling Reinforcement Learning and Decision-Making for Large-Scale Infrastructure Management5 Sep 2024 0 repositories listed
-
Prediction Accuracy & Reliability: Classification and Object Localization under Distribution Shift5 Sep 2024 0 repositories listed
-
Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression5 Sep 2024 0 repositories listed
-
NUMOSIM: A Synthetic Mobility Dataset with Anomaly Detection Benchmarks4 Sep 2024 0 repositories listed
-
PUB: Plot Understanding Benchmark and Dataset for Evaluating Large Language Models on Synthetic Visual Data Interpretation4 Sep 2024 0 repositories listed
-
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture3 Sep 2024 0 repositories listed
-
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision3 Sep 2024 0 repositories listed
-
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents3 Sep 2024 0 repositories listed
-
A practical generalization metric for deep networks benchmarking2 Sep 2024 0 repositories listed
-
Landscape-Aware Automated Algorithm Configuration using Multi-output Mixed Regression and Classification2 Sep 2024 0 repositories listed
-
Revisiting Safe Exploration in Safe Reinforcement learning2 Sep 2024 0 repositories listed
-
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages1 Sep 2024 0 repositories listed
-
Accelerating the discovery of steady-states of planetary interior dynamics with machine learning30 Aug 2024 0 repositories listed
-
Understanding the User: An Intent-Based Ranking Dataset30 Aug 2024 0 repositories listed
-
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction29 Aug 2024 0 repositories listed
-
Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games28 Aug 2024 0 repositories listed
-
Benchmarking foundation models as feature extractors for weakly-supervised computational pathology28 Aug 2024 0 repositories listed
-
Applications in CityLearn Gym Environment for Multi-Objective Control Benchmarking in Grid-Interactive Buildings and Districts27 Aug 2024 0 repositories listed
-
Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper27 Aug 2024 0 repositories listed
-
BOX3D: Lightweight Camera-LiDAR Fusion for 3D Object Detection and Localization27 Aug 2024 0 repositories listed
-
Cross-subject Brain Functional Connectivity Analysis for Multi-task Cognitive State Evaluation27 Aug 2024 0 repositories listed
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis27 Aug 2024 0 repositories listed
-
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study26 Aug 2024 0 repositories listed
-
K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences26 Aug 2024 0 repositories listed
-
DHP Benchmark: Are LLMs Good NLG Evaluators?25 Aug 2024 0 repositories listed
-
Data Augmentation for Continual RL via Adversarial Gradient Episodic Memory24 Aug 2024 0 repositories listed
-
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA24 Aug 2024 0 repositories listed
-
Open Llama2 Model for the Lithuanian Language23 Aug 2024 0 repositories listed
-
S3Simulator: A benchmarking Side Scan Sonar Simulator dataset for Underwater Image Analysis23 Aug 2024 0 repositories listed
-
Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection23 Aug 2024 0 repositories listed
-
Benchmarking Counterfactual Interpretability in Deep Learning Models for Time Series Classification22 Aug 2024 0 repositories listed
-
Dynamic PDB: A New Dataset and a SE(3) Model Extension by Integrating Dynamic Behaviors and Physical Properties in Protein Structures22 Aug 2024 0 repositories listed
-
Extraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network Analysis22 Aug 2024 0 repositories listed
-
MultiMed: Massively Multimodal and Multitask Medical Understanding22 Aug 2024 0 repositories listed
-
Advances in Preference-based Reinforcement Learning: A Review21 Aug 2024 0 repositories listed
-
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain21 Aug 2024 0 repositories listed
-
ISLES'24: Improving final infarct prediction in ischemic stroke using multimodal imaging and clinical data20 Aug 2024 0 repositories listed
-
QPO: Query-dependent Prompt Optimization via Multi-Loop Offline Reinforcement Learning20 Aug 2024 0 repositories listed
-
RP1M: A Large-Scale Motion Dataset for Piano Playing with Bi-Manual Dexterous Robot Hands20 Aug 2024 0 repositories listed
-
UKAN: Unbound Kolmogorov-Arnold Network Accompanied with Accelerated Library20 Aug 2024 0 repositories listed
-
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving19 Aug 2024 0 repositories listed
-
Benchmarking the Capabilities of Large Language Models in Transportation System Engineering: Accuracy, Consistency, and Reasoning Behaviors15 Aug 2024 0 repositories listed
-
A Meta-Engine Framework for Interleaved Task and Motion Planning using Topological Refinements11 Aug 2024 0 repositories listed
-
A Novel Momentum-Based Deep Learning Techniques for Medical Image Classification and Segmentation11 Aug 2024 0 repositories listed
-
Benchmarking Conventional and Learned Video Codecs with a Low-Delay Configuration9 Aug 2024 0 repositories listed
-
h4rm3l: A language for Composable Jailbreak Attack Synthesis9 Aug 2024 0 repositories listed
-
FedAD-Bench: A Unified Benchmark for Federated Unsupervised Anomaly Detection in Tabular Data8 Aug 2024 0 repositories listed
-
SegXAL: Explainable Active Learning for Semantic Segmentation in Driving Scene Scenarios8 Aug 2024 0 repositories listed
-
Towards Explainable Network Intrusion Detection using Large Language Models8 Aug 2024 0 repositories listed
-
Online Model-based Anomaly Detection in Multivariate Time Series: Taxonomy, Survey, Research Challenges and Future Directions7 Aug 2024 0 repositories listed
-
Soft-Hard Attention U-Net Model and Benchmark Dataset for Multiscale Image Shadow Removal7 Aug 2024 0 repositories listed
-
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline6 Aug 2024 0 repositories listed
-
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future5 Aug 2024 0 repositories listed
-
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities5 Aug 2024 0 repositories listed
-
SPINEX-TimeSeries: Similarity-based Predictions with Explainable Neighbors Exploration for Time Series and Forecasting Problems4 Aug 2024 0 repositories listed
-
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance4 Aug 2024 0 repositories listed
-
Deep Reinforcement Learning for Dynamic Order Picking in Warehouse Operations3 Aug 2024 0 repositories listed
-
IBB Traffic Graph Data: Benchmarking and Road Traffic Prediction Model2 Aug 2024 0 repositories listed
-
PINNs for Medical Image Analysis: A Survey2 Aug 2024 0 repositories listed
-
IN-Sight: Interactive Navigation through Sight1 Aug 2024 0 repositories listed
-
Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model31 Jul 2024 0 repositories listed
-
KemenkeuGPT: Leveraging a Large Language Model on Indonesia's Government Financial Data and Regulations to Enhance Decision Making31 Jul 2024 0 repositories listed
-
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models30 Jul 2024 0 repositories listed
-
Benchmarking Histopathology Foundation Models for Ovarian Cancer Bevacizumab Treatment Response Prediction from Whole Slide Images30 Jul 2024 0 repositories listed
-
Efficient Channel Estimation for Millimeter Wave and Terahertz Systems Enabled by Integrated Super-resolution Sensing and Communication30 Jul 2024 0 repositories listed
-
GNUMAP: A Parameter-Free Approach to Unsupervised Dimensionality Reduction via Graph Neural Networks30 Jul 2024 0 repositories listed
-
Anomalous State Sequence Modeling to Enhance Safety in Reinforcement Learning29 Jul 2024 0 repositories listed
-
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks29 Jul 2024 0 repositories listed
-
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection28 Jul 2024 0 repositories listed
-
On the Evaluation Consistency of Attribution-based Explanations28 Jul 2024 0 repositories listed
-
Towards a Multidimensional Evaluation Framework for Empathetic Conversational Systems26 Jul 2024 0 repositories listed
-
GermanPartiesQA: Benchmarking Commercial Large Language Models for Political Bias and Sycophancy25 Jul 2024 0 repositories listed
-
SMiCRM: A Benchmark Dataset of Mechanistic Molecular Images25 Jul 2024 0 repositories listed
-
Building a Domain-specific Guardrail Model in Production24 Jul 2024 0 repositories listed
-
Quality Assured: Rethinking Annotation Strategies in Imaging AI24 Jul 2024 0 repositories listed
-
Can time series forecasting be automated? A benchmark and analysis23 Jul 2024 0 repositories listed
-
Flexible Generation of Preference Data for Recommendation Analysis23 Jul 2024 0 repositories listed
-
Hi-EF: Benchmarking Emotion Forecasting in Human-interaction23 Jul 2024 0 repositories listed
-
Benchmarks as Microscopes: A Call for Model Metrology22 Jul 2024 0 repositories listed
-
Cascaded two-stage feature clustering and selection via separability and consistency in fuzzy decision systems22 Jul 2024 0 repositories listed