Methods › General › Attention Modules › Multi-Head Attention › Papers, page 56
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 56 of 249: papers 5,501 to 5,600 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation 26 Jul 2024 · 0 repositories · arXiv:2407.18715
-
Deep Companion Learning: Enhancing Generalization Through Historical Consistency 26 Jul 2024 · 0 repositories · arXiv:2407.18821
-
GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and Doves 26 Jul 2024 · 1 repository · arXiv:2407.19110
-
Human-artificial intelligence teaming for scientific information extraction from data-driven additive manufacturing research using large language models 26 Jul 2024 · 0 repositories · arXiv:2407.18827
-
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks 26 Jul 2024 · 1 repository · arXiv:2407.18525Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
MistralBSM: Leveraging Mistral-7B for Vehicular Networks Misbehavior Detection 26 Jul 2024 · 0 repositories · arXiv:2407.18462
-
Mixed Non-linear Quantization for Vision Transformers 26 Jul 2024 · 1 repository · arXiv:2407.18437
-
Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention 26 Jul 2024 · 1 repository · arXiv:2407.18552
-
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation 26 Jul 2024 · 1 repository · arXiv:2407.19056Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
QT-TDM: Planning With Transformer Dynamics Model and Autoregressive Q-Learning 26 Jul 2024 · 0 repositories · arXiv:2407.18841
-
REAPER: Reasoning based Retrieval Planning for Complex RAG Systems 26 Jul 2024 · 0 repositories · arXiv:2407.18553
-
SHIC: Shape-Image Correspondences with no Keypoint Supervision 26 Jul 2024 · 0 repositories · arXiv:2407.18907
-
Skin Cancer Detection utilizing Deep Learning: Classification of Skin Lesion Images using a Vision Transformer 26 Jul 2024 · 0 repositories · arXiv:2407.18554
-
TAGIFY: LLM-powered Tagging Interface for Improved Data Findability on OGD portals 26 Jul 2024 · 0 repositories · arXiv:2407.18764
-
Towards a Transformer-Based Pre-trained Model for IoT Traffic Classification 26 Jul 2024 · 1 repository · arXiv:2407.19051
-
Using GPT-4 to guide causal machine learning 26 Jul 2024 · 0 repositories · arXiv:2407.18607
-
Using Large Language Models for the Interpretation of Building Regulations 26 Jul 2024 · 0 repositories · arXiv:2407.21060
-
Adversarially Robust Decision Transformer 25 Jul 2024 · 1 repository · arXiv:2407.18414Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Banyan: Improved Representation Learning with Explicit Structure 25 Jul 2024 · 0 repositories · arXiv:2407.17771
-
Closing the gap between open-source and commercial large language models for medical evidence summarization 25 Jul 2024 · 0 repositories · arXiv:2408.00588
-
Cost-effective Instruction Learning for Pathology Vision and Language Analysis 25 Jul 2024 · 1 repository · arXiv:2407.17734
-
CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation 25 Jul 2024 · 1 repository · arXiv:2407.18070
-
Detection of manatee vocalisations using the Audio Spectrogram Transformer 25 Jul 2024 · 1 repository · arXiv:2407.18083
-
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline 25 Jul 2024 · 0 repositories · arXiv:2407.17879
-
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era? 25 Jul 2024 · 0 repositories · arXiv:2407.17870
-
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption 25 Jul 2024 · 0 repositories · arXiv:2407.18003
-
PEFT-U: Parameter-Efficient Fine-Tuning for User Personalization 25 Jul 2024 · 1 repository · arXiv:2407.18078Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
PersonaGym: Evaluating Persona Agents and LLMs 25 Jul 2024 · 1 repository · arXiv:2407.18416
-
Positive Text Reframing under Multi-strategy Optimization 25 Jul 2024 · 1 repository · arXiv:2407.17940
-
RoBERTa, ResNeXt and BiLSTM with self-attention: The ultimate trio for customer sentiment analysis 25 Jul 2024 · 0 repositories
-
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning 25 Jul 2024 · 1 repository · arXiv:2407.18248Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
The Geometry of Queries: Query-Based Innovations in Retrieval-Augmented Generation 25 Jul 2024 · 0 repositories · arXiv:2407.18044
-
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition 25 Jul 2024 · 0 repositories · arXiv:2407.18249
-
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement 25 Jul 2024 · 0 repositories · arXiv:2407.18370
-
Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT 25 Jul 2024 · 0 repositories · arXiv:2407.21058
-
What Matters in Explanations: Towards Explainable Fake Review Detection Focusing on Transformers 24 Jul 2024 · 0 repositories · arXiv:2407.21056
-
A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance 24 Jul 2024 · 0 repositories · arXiv:2407.17383
-
A Novel Two-Step Fine-Tuning Pipeline for Cold-Start Active Learning in Text Classification Tasks 24 Jul 2024 · 0 repositories · arXiv:2407.17284
-
Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications 24 Jul 2024 · 0 repositories · arXiv:2407.21055
-
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric 24 Jul 2024 · 1 repository · arXiv:2407.16981
-
Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models 24 Jul 2024 · 1 repository · arXiv:2407.17406
-
Domain Generalized Recaptured Screen Image Identification Using SWIN Transformer 24 Jul 2024 · 0 repositories · arXiv:2407.17170
-
Dynamic Graph Transformer with Correlated Spatial-Temporal Positional Encoding 24 Jul 2024 · 1 repository · arXiv:2407.16959
-
Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation 24 Jul 2024 · 1 repository · arXiv:2407.17261
-
Graph Neural Networks: A suitable Alternative to MLPs in Latent 3D Medical Image Classification? 24 Jul 2024 · 1 repository · arXiv:2407.17219
-
I Could've Asked That: Reformulating Unanswerable Questions 24 Jul 2024 · 1 repository · arXiv:2407.17469
-
Improving ICD coding using Chapter based Named Entities and Attentional Models 24 Jul 2024 · 0 repositories · arXiv:2407.17230
-
LoFormer: Local Frequency Transformer for Image Deblurring 24 Jul 2024 · 2 repositories · arXiv:2407.16993
-
MuST: Multi-Scale Transformers for Surgical Phase Recognition 24 Jul 2024 · 1 repository · arXiv:2407.17361
-
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge 24 Jul 2024 · 0 repositories · arXiv:2408.01453
-
Testing Large Language Models on Driving Theory Knowledge and Skills for Connected Autonomous Vehicles 24 Jul 2024 · 0 repositories · arXiv:2407.17211
-
Trans2Unet: Neural fusion for Nuclei Semantic Segmentation 24 Jul 2024 · 0 repositories · arXiv:2407.17181
-
Analyzing Polysemy Evolution Using Semantic Cells 23 Jul 2024 · 0 repositories · arXiv:2407.16110
-
Artificial Intelligence in Extracting Diagnostic Data from Dental Records 23 Jul 2024 · 0 repositories · arXiv:2407.21050
-
Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution 23 Jul 2024 · 0 repositories · arXiv:2407.16232
-
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data? 23 Jul 2024 · 1 repository · arXiv:2407.16607Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Diffusion Transformer Captures Spatial-Temporal Dependencies: A Theory for Gaussian Process Data 23 Jul 2024 · 0 repositories · arXiv:2407.16134
-
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models 23 Jul 2024 · 0 repositories · arXiv:2407.16221
-
Enhancing LLM's Cognition via Structurization 23 Jul 2024 · 1 repository · arXiv:2407.16434Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Exploring The Neural Burden In Pruned Models: An Insight Inspired By Neuroscience 23 Jul 2024 · 0 repositories · arXiv:2407.16716
-
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification 23 Jul 2024 · 0 repositories · arXiv:2407.16244
-
HyTAS: A Hyperspectral Image Transformer Architecture Search Benchmark and Analysis 23 Jul 2024 · 1 repository · arXiv:2407.16269
-
LawLuo: A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation 23 Jul 2024 · 0 repositories · arXiv:2407.16252
-
Lawma: The Power of Specialization for Legal Tasks 23 Jul 2024 · 0 repositories · arXiv:2407.16615Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Masked Graph Learning with Recurrent Alignment for Multimodal Emotion Recognition in Conversation 23 Jul 2024 · 0 repositories · arXiv:2407.16714
-
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection 23 Jul 2024 · 1 repository · arXiv:2407.16237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Patched RTC: evaluating LLMs for diverse software development tasks 23 Jul 2024 · 1 repository · arXiv:2407.16557
-
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent 23 Jul 2024 · 0 repositories · arXiv:2407.16667
-
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach 23 Jul 2024 · 0 repositories · arXiv:2407.16833
-
Robust Privacy Amidst Innovation with Large Language Models Through a Critical Assessment of the Risks 23 Jul 2024 · 1 repository · arXiv:2407.16166
-
S-E Pipeline: A Vision Transformer (ViT) based Resilient Classification Pipeline for Medical Imaging Against Adversarial Attacks 23 Jul 2024 · 0 repositories · arXiv:2407.17587
-
SINDER: Repairing the Singular Defects of DINOv2 23 Jul 2024 · 1 repository · arXiv:2407.16826Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Synthesizer Sound Matching Using Audio Spectrogram Transformers 23 Jul 2024 · 0 repositories · arXiv:2407.16643
-
TookaBERT: A Step Forward for Persian NLU 23 Jul 2024 · 0 repositories · arXiv:2407.16382
-
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval 22 Jul 2024 · 0 repositories · arXiv:2408.03340
-
Can GPT-4 learn to analyse moves in research article abstracts? 22 Jul 2024 · 0 repositories · arXiv:2407.15612
-
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA 22 Jul 2024 · 0 repositories · arXiv:2407.15353
-
Dissecting Multiplication in Transformers: Insights into LLMs 22 Jul 2024 · 1 repository · arXiv:2407.15360
-
Efficient Multi-disparity Transformer for Light Field Image Super-resolution 22 Jul 2024 · 0 repositories · arXiv:2407.15329
-
Estimating Probability Densities with Transformer and Denoising Diffusion 22 Jul 2024 · 1 repository · arXiv:2407.15703Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Impacts of Anthropomorphizing Large Language Models in Learning Environments 22 Jul 2024 · 0 repositories · arXiv:2408.03945
-
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models 22 Jul 2024 · 0 repositories · arXiv:2407.15399
-
Inverted Activations: Reducing Memory Footprint in Neural Network Training 22 Jul 2024 · 1 repository · arXiv:2407.15545
-
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer 22 Jul 2024 · 0 repositories · arXiv:2407.16026
-
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning 22 Jul 2024 · 0 repositories · arXiv:2407.15815
-
LLMmap: Fingerprinting For Large Language Models 22 Jul 2024 · 1 repository · arXiv:2407.15847Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 16 harvested samples)
-
Local All-Pair Correspondence for Point Tracking 22 Jul 2024 · 2 repositories · arXiv:2407.15420Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training 22 Jul 2024 · 1 repository · arXiv:2407.15892
-
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity 22 Jul 2024 · 1 repository · arXiv:2407.15838Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
MoRSE: Bridging the Gap in Cybersecurity Expertise with Retrieval Augmented Generation 22 Jul 2024 · 0 repositories · arXiv:2407.15748
-
NV-Retriever: Improving text embedding models with effective hard-negative mining 22 Jul 2024 · 0 repositories · arXiv:2407.15831
-
Predicting the Best of N Visual Trackers 22 Jul 2024 · 1 repository · arXiv:2407.15707
-
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines 22 Jul 2024 · 1 repository · arXiv:2407.21046
-
RadioRAG: Factual large language models for enhanced diagnostics in radiology using online retrieval augmented generation 22 Jul 2024 · 1 repository · arXiv:2407.15621
-
Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget 22 Jul 2024 · 1 repository · arXiv:2407.15811Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Unlocking the Potential: Benchmarking Large Language Models in Water Engineering and Research 22 Jul 2024 · 0 repositories · arXiv:2407.21045
-
ZZU-NLP at SIGHAN-2024 dimABSA Task: Aspect-Based Sentiment Analysis with Coarse-to-Fine In-context Learning 22 Jul 2024 · 0 repositories · arXiv:2407.15341
-
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts 21 Jul 2024 · 0 repositories · arXiv:2407.15136
-
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts 21 Jul 2024 · 0 repositories · arXiv:2407.15050
-
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment 21 Jul 2024 · 1 repository · arXiv:2407.15184