Browse State-of-the-Art › Benchmarking › Papers, page 43
Benchmarking
Papers archive 2025-07-28
archive papers tagged: 5,548 · with a code link: 2,658 · where Syntology ran a sample: 749 (624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (749 of 5,548 tagged: 624 with a run with no instrument failure, 125 where every run was a failure of Syntology's instrument)
Page 43 of 56: papers 4,201 to 4,300 of 5,548, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Optimizing with Low Budgets: a Comparison on the Black-box Optimization Benchmarking Suite and OpenAI Gym29 Sep 2023 0 repositories listed
-
Sarcasm in Sight and Sound: Benchmarking and Expansion to Improve Multimodal Sarcasm Detection29 Sep 2023 0 repositories listed
-
Language Models as a Service: Overview of a New Paradigm and its Challenges28 Sep 2023 0 repositories listed
-
Demographic Parity: Mitigating Biases in Real-World Data27 Sep 2023 0 repositories listed
-
Advancing The Rate-Distortion-Computation Frontier For Neural Image Compression26 Sep 2023 0 repositories listed
-
On quantifying and improving realism of images generated with diffusion26 Sep 2023 0 repositories listed
-
Optimization Techniques for a Physical Model of Human Vocalisation26 Sep 2023 0 repositories listed
-
Thalamic nuclei segmentation from T₁-weighted MRI: unifying and benchmarking state-of-the-art methods with young and old cohorts26 Sep 2023 0 repositories listed
-
Efficient Pauli channel estimation with logarithmic quantum memory25 Sep 2023 0 repositories listed
-
Categorization and analysis of 14 computational methods for estimating cell potency from single-cell RNA-seq data24 Sep 2023 0 repositories listed
-
VisionKG: Unleashing the Power of Visual Datasets via Knowledge Graph24 Sep 2023 0 repositories listed
-
Domain Adaptation for Arabic Machine Translation: The Case of Financial Texts22 Sep 2023 0 repositories listed
-
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam21 Sep 2023 0 repositories listed
-
Multimodal Deep Learning for Scientific Imaging Interpretation21 Sep 2023 0 repositories listed
-
On the relationship between Benchmarking, Standards and Certification in Robotics and AI21 Sep 2023 0 repositories listed
-
Towards Effective Disambiguation for Machine Translation with Large Language Models20 Sep 2023 0 repositories listed
-
SHOWMe: Benchmarking Object-agnostic Hand-Object 3D Reconstruction19 Sep 2023 0 repositories listed
-
Training neural mapping schemes for satellite altimetry with simulation data19 Sep 2023 0 repositories listed
-
The Protein Engineering Tournament: An Open Science Benchmark for Protein Modeling and Design18 Sep 2023 0 repositories listed
-
Emerging Approaches for THz Array Imaging: A Tutorial Review and Software Tool16 Sep 2023 0 repositories listed
-
Exploration of TPUs for AI Applications16 Sep 2023 0 repositories listed
-
Benchmarking machine learning models for quantum state classification14 Sep 2023 0 repositories listed
-
Leveraging Contextual Information for Effective Entity Salience Detection14 Sep 2023 0 repositories listed
-
So you think you can track?13 Sep 2023 0 repositories listed
-
AmodalSynthDrive: A Synthetic Amodal Perception Dataset for Autonomous Driving12 Sep 2023 0 repositories listed
-
Unveiling the potential of large language models in generating semantic and cross-language clones12 Sep 2023 0 repositories listed
-
Better Practices for Domain Adaptation7 Sep 2023 0 repositories listed
-
DBsurf: A Discrepancy Based Method for Discrete Stochastic Gradient Estimation7 Sep 2023 0 repositories listed
-
Are SNNs Truly Energy-efficient? - A Hardware Perspective6 Sep 2023 0 repositories listed
-
Neural Networks for Fast Optimisation in Model Predictive Control: A Review6 Sep 2023 0 repositories listed
-
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking5 Sep 2023 0 repositories listed
-
AGIBench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark for Large Language Models5 Sep 2023 0 repositories listed
-
Hybrid data driven/thermal simulation model for comfort assessment4 Sep 2023 0 repositories listed
-
3 Sep 2023 0 repositories listed
-
Holistic Dynamic Frequency Transformer for Image Fusion and Exposure Correction3 Sep 2023 0 repositories listed
-
Can humans help BERT gain "confidence"?31 Aug 2023 0 repositories listed
-
Benchmarking Robustness and Generalization in Multi-Agent Systems: A Case Study on Neural MMO30 Aug 2023 0 repositories listed
-
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads28 Aug 2023 0 repositories listed
-
Benchmarking Data Efficiency and Computational Efficiency of Temporal Action Localization Models24 Aug 2023 0 repositories listed
-
Benchmarking Causal Study to Interpret Large Language Models for Source Code23 Aug 2023 0 repositories listed
-
Efficient Benchmarking of Language Models22 Aug 2023 0 repositories listed
-
Measuring the Effect of Causal Disentanglement on the Adversarial Robustness of Neural Network Models21 Aug 2023 0 repositories listed
-
Benchmarking Adversarial Robustness of Compressed Deep Learning Models16 Aug 2023 0 repositories listed
-
A Survey on Model Compression for Large Language Models15 Aug 2023 0 repositories listed
-
Deep Neural Operator Driven Real Time Inference for Nuclear Systems to Enable Digital Twin Solutions15 Aug 2023 0 repositories listed
-
Does AI for science need another ImageNet Or totally different benchmarks? A case study of machine learning force fields11 Aug 2023 0 repositories listed
-
Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human Evaluation10 Aug 2023 0 repositories listed
-
Spintronics for image recognition: performance benchmarking via ultrafast data-driven simulations10 Aug 2023 0 repositories listed
-
Enhancing Architecture Frameworks by Including Modern Stakeholders and their Views/Viewpoints9 Aug 2023 0 repositories listed
-
Benchmarking LLM powered Chatbots: Methods and Metrics8 Aug 2023 0 repositories listed
-
RECipe: Does a Multi-Modal Recipe Knowledge Graph Fit a Multi-Purpose Recommendation System?8 Aug 2023 0 repositories listed
-
Microvasculature Segmentation in Human BioMolecular Atlas Program (HuBMAP)6 Aug 2023 0 repositories listed
-
A Survey of Spanish Clinical Language Models4 Aug 2023 0 repositories listed
-
RobustMQ: Benchmarking Robustness of Quantized Models4 Aug 2023 0 repositories listed
-
Benchmarking Adaptative Variational Quantum Algorithms on QUBO Instances3 Aug 2023 0 repositories listed
-
Capsa: A Unified Framework for Quantifying Risk in Deep Neural Networks1 Aug 2023 0 repositories listed
-
CLAMS: A Cluster Ambiguity Measure for Estimating Perceptual Variability in Visual Clustering1 Aug 2023 0 repositories listed
-
Differential Privacy for Adaptive Weight Aggregation in Federated Tumor Segmentation1 Aug 2023 0 repositories listed
-
Deep Learning and Computer Vision for Glaucoma Detection: A Review31 Jul 2023 0 repositories listed
-
Benchmarking Performance of Deep Learning Model for Material Segmentation on Two HPC Systems27 Jul 2023 0 repositories listed
-
Fluorescent Neuronal Cells v2: Multi-Task, Multi-Format Annotations for Deep Learning in Microscopy26 Jul 2023 0 repositories listed
-
YOLOBench: Benchmarking Efficient Object Detectors on Embedded Systems26 Jul 2023 0 repositories listed
-
Benchmarking and Analyzing Generative Data for Visual Recognition25 Jul 2023 0 repositories listed
-
Implementing and Benchmarking the Locally Competitive Algorithm on the Loihi 2 Neuromorphic Processor25 Jul 2023 0 repositories listed
-
Towards an AI Accountability Policy25 Jul 2023 0 repositories listed
-
Towards Long-Term predictions of Turbulence using Neural Operators25 Jul 2023 0 repositories listed
-
UPREVE: An End-to-End Causal Discovery Benchmarking System25 Jul 2023 0 repositories listed
-
The Impact of Genomic Variation on Function (IGVF) Consortium24 Jul 2023 0 repositories listed
-
The Extractive-Abstractive Axis: Measuring Content "Borrowing" in Generative Language Models20 Jul 2023 0 repositories listed
-
Approaches for benchmarking single-cell gene regulatory network inference methods17 Jul 2023 0 repositories listed
-
Benchmarking fixed-length Fingerprint Representations across different Embedding Sizes and Sensor Types17 Jul 2023 0 repositories listed
-
Machine Learning for Ranking f-wave Extraction Methods in Single-Lead ECGs17 Jul 2023 0 repositories listed
-
On the Real-Time Semantic Segmentation of Aphid Clusters in the Wild17 Jul 2023 0 repositories listed
-
Revisiting Implicit Models: Sparsity Trade-offs Capability in Weight-tied Model for Vision Tasks16 Jul 2023 0 repositories listed
-
Benchmarking the Effectiveness of Classification Algorithms and SVM Kernels for Dry Beans15 Jul 2023 0 repositories listed
-
Joint Batching and Scheduling for High-Throughput Multiuser Edge AI with Asynchronous Task Arrivals15 Jul 2023 0 repositories listed
-
Benchmarking Explanatory Models for Inertia Forecasting using Public Data of the Nordic Area14 Jul 2023 0 repositories listed
-
Challenge Results Are Not Reproducible14 Jul 2023 0 repositories listed
-
Deep Generative Models for Physiological Signals: A Systematic Literature Review12 Jul 2023 0 repositories listed
-
Pathway: a fast and flexible unified stream data processing framework for analytical and Machine Learning applications12 Jul 2023 0 repositories listed
-
Benchmarking Bayesian Causal Discovery Methods for Downstream Treatment Effect Estimation11 Jul 2023 0 repositories listed
-
Temporal Graphs Anomaly Emergence Detection: Benchmarking For Social Media Interactions11 Jul 2023 0 repositories listed
-
Assessing the efficacy of large language models in generating accurate teacher responses9 Jul 2023 0 repositories listed
-
Fairness-Aware Graph Neural Networks: A Survey8 Jul 2023 0 repositories listed
-
Fast Empirical Scenarios8 Jul 2023 0 repositories listed
-
Structural Property Prediction5 Jul 2023 0 repositories listed
-
Unsupervised Spectral Demosaicing with Lightweight Spectral Attention Networks5 Jul 2023 0 repositories listed
-
OpenSiteRec: An Open Dataset for Site Recommendation3 Jul 2023 0 repositories listed
-
A Synthetic Benchmarking Pipeline to Compare Camera Calibration Algorithms3 Jul 2023 0 repositories listed
-
Conditionally Invariant Representation Learning for Disentangling Cellular Heterogeneity2 Jul 2023 0 repositories listed
-
InstructEval: Systematic Evaluation of Instruction Selection Methods1 Jul 2023 0 repositories listed
-
SysNoise: Exploring and Benchmarking Training-Deployment System Inconsistency1 Jul 2023 0 repositories listed
-
Benchmarking Large Language Model Capabilities for Conditional Generation29 Jun 2023 0 repositories listed
-
Generative AI for Programming Education: Benchmarking ChatGPT, GPT-4, and Human Tutors29 Jun 2023 0 repositories listed
-
Learning Environment Models with Continuous Stochastic Dynamics29 Jun 2023 0 repositories listed
-
Principles and Guidelines for Evaluating Social Robot Navigation Algorithms29 Jun 2023 0 repositories listed
-
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity28 Jun 2023 0 repositories listed
-
Effective Transfer of Pretrained Large Visual Model for Fabric Defect Segmentation via Specifc Knowledge Injection28 Jun 2023 0 repositories listed
-
Emotion Analysis of Tweets Banning Education in Afghanistan28 Jun 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.