Browse State-of-the-Art › Benchmarking › Papers, page 36
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 36 of 56: papers 3,501 to 3,600 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
FuzzWiz -- Fuzzing Framework for Efficient Hardware Coverage23 Oct 2024 0 repositories listed
-
Benchmarking Smoothness and Reducing High-Frequency Oscillations in Continuous Control Policies22 Oct 2024 0 repositories listed
-
Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp Editing22 Oct 2024 0 repositories listed
-
Safe Load Balancing in Software-Defined-Networking22 Oct 2024 0 repositories listed
-
A Framework for Evaluating Predictive Models Using Synthetic Image Covariates and Longitudinal Data21 Oct 2024 0 repositories listed
-
Hiding in Plain Sight: Reframing Hardware Trojan Benchmarking as a Hide&Seek Modification21 Oct 2024 0 repositories listed
-
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping21 Oct 2024 0 repositories listed
-
Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence20 Oct 2024 0 repositories listed
-
Advancing Histopathology with Deep Learning Under Data Scarcity: A Decade in Review18 Oct 2024 0 repositories listed
-
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs18 Oct 2024 0 repositories listed
-
Sum Secrecy Rate Maximization for Full Duplex ISAC Systems17 Oct 2024 0 repositories listed
-
Trust but Verify: Programmatic VLM Evaluation in the Wild17 Oct 2024 0 repositories listed
-
AERO: Softmax-Only LLMs for Efficient Private Inference16 Oct 2024 0 repositories listed
-
Benchmarking Defeasible Reasoning with Large Language Models -- Initial Experiments and Future Directions16 Oct 2024 0 repositories listed
-
Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation16 Oct 2024 0 repositories listed
-
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs16 Oct 2024 0 repositories listed
-
Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos15 Oct 2024 0 repositories listed
-
FoundTS: Comprehensive and Unified Benchmarking of Foundation Models for Time Series Forecasting15 Oct 2024 0 repositories listed
-
Building a Multivariate Time Series Benchmarking Datasets Inspired by Natural Language Processing (NLP)14 Oct 2024 0 repositories listed
-
ChakmaNMT: A Low-resource Machine Translation On Chakma Language14 Oct 2024 0 repositories listed
-
Personalised Feedback Framework for Online Education Programmes Using Generative AI14 Oct 2024 0 repositories listed
-
The Trap of Presumed Equivalence: Artificial General Intelligence Should Not Be Assessed on the Scale of Human Intelligence14 Oct 2024 0 repositories listed
-
Transforming Game Play: A Comparative Study of DCQN and DTQN Architectures in Reinforcement Learning14 Oct 2024 0 repositories listed
-
A Comparative Analysis on Ethical Benchmarking in Large Language Models11 Oct 2024 0 repositories listed
-
Can we hop in general? A discussion of benchmark selection and design using the Hopper environment11 Oct 2024 0 repositories listed
-
∀uto∃∨∧L: Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks11 Oct 2024 0 repositories listed
-
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation11 Oct 2024 0 repositories listed
-
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example11 Oct 2024 0 repositories listed
-
Advocating Character Error Rate for Multilingual ASR Evaluation9 Oct 2024 0 repositories listed
-
Analysis of different disparity estimation techniques on aerial stereo image datasets9 Oct 2024 0 repositories listed
-
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding9 Oct 2024 0 repositories listed
-
InAttention: Linear Context Scaling for Transformers9 Oct 2024 0 repositories listed
-
M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes9 Oct 2024 0 repositories listed
-
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB9 Oct 2024 0 repositories listed
-
Active Evaluation Acquisition for Efficient LLM Benchmarking8 Oct 2024 0 repositories listed
-
Benchmarking of a new data splitting method on volcanic eruption data8 Oct 2024 0 repositories listed
-
Manual Verbalizer Enrichment for Few-Shot Text Classification8 Oct 2024 0 repositories listed
-
Precise Model Benchmarking with Only a Few Observations7 Oct 2024 0 repositories listed
-
Rule-based Data Selection for Large Language Models7 Oct 2024 0 repositories listed
-
Translation Canvas: An Explainable Interface to Pinpoint and Analyze Translation Systems7 Oct 2024 0 repositories listed
-
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection6 Oct 2024 0 repositories listed
-
Implicit to Explicit Entropy Regularization: Benchmarking ViT Fine-tuning under Noisy Labels5 Oct 2024 0 repositories listed
-
PalmBench: A Comprehensive Benchmark of Compressed Large Language Models on Mobile Platforms5 Oct 2024 0 repositories listed
-
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends5 Oct 2024 0 repositories listed
-
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities4 Oct 2024 0 repositories listed
-
4 Oct 2024 0 repositories listed Syntology 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Understanding Large Language Models in Your Pockets: Performance Study on COTS Mobile Devices4 Oct 2024 0 repositories listed
-
PersoBench: Benchmarking Personalized Response Generation in Large Language Models4 Oct 2024 0 repositories listed
-
IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models3 Oct 2024 0 repositories listed
-
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning3 Oct 2024 0 repositories listed
-
Repurposing Foundation Model for Generalizable Medical Time Series Classification3 Oct 2024 0 repositories listed
-
A Real Benchmark Swell Noise Dataset for Performing Seismic Data Denoising via Deep Learning2 Oct 2024 0 repositories listed
-
CALF: Benchmarking Evaluation of LFQA Using Chinese Examinations2 Oct 2024 0 repositories listed
-
ConServe: Harvesting GPUs for Low-Latency and High-Throughput Large Language Model Serving2 Oct 2024 0 repositories listed
-
Deep learning for action spotting in association football videos2 Oct 2024 0 repositories listed
-
Deep Unlearn: Benchmarking Machine Unlearning2 Oct 2024 0 repositories listed
-
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description2 Oct 2024 0 repositories listed
-
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs2 Oct 2024 0 repositories listed
-
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents1 Oct 2024 0 repositories listed
-
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks1 Oct 2024 0 repositories listed
-
Benchmarking Adaptive Intelligence and Computer Vision on Human-Robot Collaboration30 Sep 2024 0 repositories listed
-
Match Stereo Videos via Bidirectional Alignment30 Sep 2024 0 repositories listed
-
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs30 Sep 2024 0 repositories listed
-
AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy29 Sep 2024 0 repositories listed
-
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks29 Sep 2024 0 repositories listed
-
Tracking Everything in Robotic-Assisted Surgery29 Sep 2024 0 repositories listed
-
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement28 Sep 2024 0 repositories listed
-
bnRep: A repository of Bayesian networks from the academic literature27 Sep 2024 0 repositories listed
-
CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting27 Sep 2024 0 repositories listed
-
Data Analysis in the Era of Generative AI27 Sep 2024 0 repositories listed
-
EarthquakeNPP: Benchmark Datasets for Earthquake Forecasting with Neural Point Processes27 Sep 2024 0 repositories listed
-
MCUBench: A Benchmark of Tiny Object Detectors on MCUs27 Sep 2024 0 repositories listed
-
Benchmarking Deep Learning Models for Object Detection on Edge Computing Devices25 Sep 2024 0 repositories listed
-
Omnibenchmark (alpha) for continuous and open benchmarking in bioinformatics25 Sep 2024 0 repositories listed
-
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning25 Sep 2024 0 repositories listed
-
SEN12-WATER: A New Dataset for Hydrological Applications and its Benchmarking25 Sep 2024 0 repositories listed
-
HLB: Benchmarking LLMs' Humanlikeness in Language Use24 Sep 2024 0 repositories listed
-
Qualitative Insights Tool (QualIT): LLM Enhanced Topic Modeling24 Sep 2024 0 repositories listed
-
Benchmarking Edge AI Platforms for High-Performance ML Inference23 Sep 2024 0 repositories listed
-
Building a continuous benchmarking ecosystem in bioinformatics23 Sep 2024 0 repositories listed
-
Sketch 'n Solve: An Efficient Python Package for Large-Scale Least Squares Using Randomized Numerical Linear Algebra22 Sep 2024 0 repositories listed
-
The Ability of Large Language Models to Evaluate Constraint-satisfaction in Agent Responses to Open-ended Requests22 Sep 2024 0 repositories listed
-
An Evolutionary Algorithm For the Vehicle Routing Problem with Drones with Interceptions21 Sep 2024 0 repositories listed
-
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology21 Sep 2024 0 repositories listed
-
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data20 Sep 2024 0 repositories listed
-
Robust Salient Object Detection on Compressed Images Using Convolutional Neural Networks20 Sep 2024 0 repositories listed
-
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection20 Sep 2024 0 repositories listed
-
Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time20 Sep 2024 0 repositories listed
-
Arena 4.0: A Comprehensive ROS2 Development and Benchmarking Platform for Human-centric Navigation Using Generative-Model-based Environment Generation19 Sep 2024 0 repositories listed
-
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines19 Sep 2024 0 repositories listed
-
Efficacy of Synthetic Data as a Benchmark18 Sep 2024 0 repositories listed
-
Quantum Kernel Learning for Small Dataset Modeling in Semiconductor Fabrication: Application to Ohmic Contact17 Sep 2024 0 repositories listed
-
WER We Stand: Benchmarking Urdu ASR Models17 Sep 2024 0 repositories listed
-
Benchmarking VLMs' Reasoning About Persuasive Atypical Images16 Sep 2024 0 repositories listed
-
Benchmarking LLMs in Political Content Text-Annotation: Proof-of-Concept with Toxicity and Incivility Data15 Sep 2024 0 repositories listed
-
Byzantine-Robust and Communication-Efficient Distributed Learning via Compressed Momentum Filtering13 Sep 2024 0 repositories listed
-
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study13 Sep 2024 0 repositories listed
-
Text-To-Speech Synthesis In The Wild13 Sep 2024 0 repositories listed
-
Efficient Sparse Coding with the Adaptive Locally Competitive Algorithm for Speech Classification12 Sep 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.