Methods › General › Attention Modules › Multi-Head Attention › Papers, page 91
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 91 of 249: papers 9,001 to 9,100 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Memory Consolidation Enables Long-Context Video Understanding 8 Feb 2024 · 0 repositories · arXiv:2402.05861
-
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data 8 Feb 2024 · 0 repositories · arXiv:2402.05545
-
Neural Models for Source Code Synthesis and Completion 8 Feb 2024 · 0 repositories · arXiv:2402.06690
-
Noise Contrastive Alignment of Language Models with Explicit Rewards 8 Feb 2024 · 3 repositories · arXiv:2402.05369Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On Convolutional Vision Transformers for Yield Prediction 8 Feb 2024 · 0 repositories · arXiv:2402.05557
-
Question Aware Vision Transformer for Multimodal Reasoning 8 Feb 2024 · 0 repositories · arXiv:2402.05472
-
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation 8 Feb 2024 · 0 repositories · arXiv:2402.05699
-
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting 8 Feb 2024 · 0 repositories · arXiv:2402.05830
-
TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation 8 Feb 2024 · 0 repositories · arXiv:2402.05733
-
Unleashing the Infinity Power of Geometry: A Novel Geometry-Aware Transformer (GOAT) for Whole Slide Histopathology Image Analysis 8 Feb 2024 · 0 repositories · arXiv:2402.05373
-
You Only Need One Color Space: An Efficient Network for Low-light Image Enhancement 8 Feb 2024 · 1 repository · arXiv:2402.05809
-
Zero-Shot Chain-of-Thought Reasoning Guided by Evolutionary Algorithms in Large Language Models 8 Feb 2024 · 0 repositories · arXiv:2402.05376
-
A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models 7 Feb 2024 · 1 repository · arXiv:2402.04787
-
Aspect-Based Sentiment Analysis for Open-Ended HR Survey Responses 7 Feb 2024 · 0 repositories · arXiv:2402.04812
-
Can Large Language Model Agents Simulate Human Trust Behavior? 7 Feb 2024 · 1 repository · arXiv:2402.04559Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Conversational Assistants in Knowledge-Intensive Contexts: An Evaluation of LLM- versus Intent-based Systems 7 Feb 2024 · 0 repositories · arXiv:2402.04955
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 · 1 repository · arXiv:2402.04494Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models 7 Feb 2024 · 0 repositories · arXiv:2402.05034
-
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach 7 Feb 2024 · 0 repositories · arXiv:2402.04609
-
Latent Plan Transformer for Trajectory Abstraction: Planning as Latent Space Inference 7 Feb 2024 · 1 repository · arXiv:2402.04647Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 2 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning 7 Feb 2024 · 1 repository · arXiv:2402.04833Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Navigating the Knowledge Sea: Planet-scale answer retrieval using LLMs 7 Feb 2024 · 0 repositories · arXiv:2402.05318
-
Opening the AI black box: program synthesis via mechanistic interpretability 7 Feb 2024 · 1 repository · arXiv:2402.05110Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training 7 Feb 2024 · 0 repositories · arXiv:2402.05033
-
StableMask: Refining Causal Masking in Decoder-only Transformer 7 Feb 2024 · 0 repositories · arXiv:2402.04779
-
Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and Calibration 7 Feb 2024 · 0 repositories · arXiv:2402.04883
-
TransLLaMa: LLM-based Simultaneous Translation System 7 Feb 2024 · 1 repository · arXiv:2402.04636
-
Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers 7 Feb 2024 · 3 repositories · arXiv:2402.04538Syntology official (archive's flag): 9 ran · 10 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Two Trades is not Baffled: Condensing Graph via Crafting Rational Gradient Matching 7 Feb 2024 · 1 repository · arXiv:2402.04924Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Advancing Legal Reasoning: The Integration of AI to Navigate Complexities and Biases in Global Jurisprudence with Semi-Automated Arbitration Processes (SAAPs) 6 Feb 2024 · 0 repositories · arXiv:2402.04140
-
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls 6 Feb 2024 · 1 repository · arXiv:2402.04253Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Attention-based Shape and Gait Representations Learning for Video-based Cloth-Changing Person Re-Identification 6 Feb 2024 · 0 repositories · arXiv:2402.03716
-
Behind the Screen: Investigating ChatGPT's Dark Personality Traits and Conspiracy Beliefs 6 Feb 2024 · 0 repositories · arXiv:2402.04110
-
Breaking Symmetry When Training Transformers 6 Feb 2024 · 0 repositories · arXiv:2402.05969
-
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks 6 Feb 2024 · 2 repositories · arXiv:2402.04248Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples)
-
CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers 6 Feb 2024 · 0 repositories · arXiv:2402.04239
-
CEHR-GPT: Generating Electronic Health Records with Chronological Patient Timelines 6 Feb 2024 · 0 repositories · arXiv:2402.04400
-
Comparing Abstraction in Humans and Large Language Models Using Multimodal Serial Reproduction 6 Feb 2024 · 0 repositories · arXiv:2402.03618
-
Detecting Mode Collapse in Language Models via Narration 6 Feb 2024 · 0 repositories · arXiv:2402.04477
-
Detection Transformer for Teeth Detection, Segmentation, and Numbering in Oral Rare Diseases: Focus on Data Augmentation and Inpainting Techniques 6 Feb 2024 · 0 repositories · arXiv:2402.04408
-
Transformer based Endmember Fusion with Spatial Context for Hyperspectral Unmixing 6 Feb 2024 · 0 repositories · arXiv:2402.03835
-
Enhancing Retrieval Processes for Language Generation with Augmented Queries 6 Feb 2024 · 0 repositories · arXiv:2402.16874
-
Identifying Reasons for Contraceptive Switching from Real-World Data Using Large Language Models 6 Feb 2024 · 1 repository · arXiv:2402.03597
-
Iterative Prompt Refinement for Radiation Oncology Symptom Extraction Using Teacher-Student Large Language Models 6 Feb 2024 · 0 repositories · arXiv:2402.04075
-
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning 6 Feb 2024 · 0 repositories · arXiv:2402.03667
-
Large Language Models As MOOCs Graders 6 Feb 2024 · 0 repositories · arXiv:2402.03776
-
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs 6 Feb 2024 · 0 repositories · arXiv:2402.03927
-
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text 6 Feb 2024 · 1 repository · arXiv:2402.04335
-
Lens: A Foundation Model for Network Traffic 6 Feb 2024 · 0 repositories · arXiv:2402.03646
-
LLM Agents can Autonomously Hack Websites 6 Feb 2024 · 0 repositories · arXiv:2402.06664
-
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification 6 Feb 2024 · 0 repositories · arXiv:2402.03686
-
Parameter-tuning-free data entry error unlearning with adaptive selective synaptic dampening 6 Feb 2024 · 1 repository · arXiv:2402.10098
-
Pard: Permutation-Invariant Autoregressive Diffusion for Graph Generation 6 Feb 2024 · 1 repository · arXiv:2402.03687
-
Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images 6 Feb 2024 · 0 repositories · arXiv:2402.03752
-
Professional Agents -- Evolving Large Language Models into Autonomous Experts with Human-Level Competencies 6 Feb 2024 · 0 repositories · arXiv:2402.03628
-
Provably learning a multi-head attention layer 6 Feb 2024 · 0 repositories · arXiv:2402.04084
-
Reinforcement Learning from Bagged Reward 6 Feb 2024 · 0 repositories · arXiv:2402.03771
-
Return-Aligned Decision Transformer 6 Feb 2024 · 0 repositories · arXiv:2402.03923
-
Self-Discover: Large Language Models Self-Compose Reasoning Structures 6 Feb 2024 · 3 repositories · arXiv:2402.03620Syntology 0 ran · 3 unverified (of 3 harvested samples)
-
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue 6 Feb 2024 · 1 repository · arXiv:2402.03658
-
Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation 6 Feb 2024 · 0 repositories · arXiv:2402.03642
-
NeRCC: Nested-Regression Coded Computing for Resilient Distributed Prediction Serving Systems 6 Feb 2024 · 0 repositories · arXiv:2402.04377
-
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry 6 Feb 2024 · 1 repository · arXiv:2402.04347
-
The Use of a Large Language Model for Cyberbullying Detection 6 Feb 2024 · 0 repositories · arXiv:2402.04088
-
Training Language Models to Generate Text with Citations via Fine-grained Rewards 6 Feb 2024 · 1 repository · arXiv:2402.04315Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
U-shaped Vision Mamba for Single Image Dehazing 6 Feb 2024 · 1 repository · arXiv:2402.04139
-
A Survey on Transformer Compression 5 Feb 2024 · 0 repositories · arXiv:2402.05964
-
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings 5 Feb 2024 · 1 repository · arXiv:2402.03172Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Arabic Synonym BERT-based Adversarial Examples for Text Classification 5 Feb 2024 · 1 repository · arXiv:2402.03477
-
C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03181Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models 5 Feb 2024 · 1 repository · arXiv:2402.02987
-
Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector 5 Feb 2024 · 2 repositories · arXiv:2402.03094Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 Feb 2024 · 5 repositories · arXiv:2402.03300Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 3 pointer-only (licence)
-
DiffsFormer: A Diffusion Transformer on Stock Factor Augmentation 5 Feb 2024 · 0 repositories · arXiv:2402.06656
-
Enhancing textual textbook question answering with large language models and retrieval augmented generation 5 Feb 2024 · 1 repository · arXiv:2402.05128Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Financial Report Chunking for Effective Retrieval Augmented Generation 5 Feb 2024 · 1 repository · arXiv:2402.05131
-
Focal Modulation Networks for Interpretable Sound Classification 5 Feb 2024 · 0 repositories · arXiv:2402.02754
-
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning 5 Feb 2024 · 1 repository · arXiv:2402.02805Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
HAMLET: Graph Transformer Neural Operator for Partial Differential Equations 5 Feb 2024 · 0 repositories · arXiv:2402.03541
-
Harnessing PubMed User Query Logs for Post Hoc Explanations of Recommended Similar Articles 5 Feb 2024 · 0 repositories · arXiv:2402.03484
-
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering 5 Feb 2024 · 0 repositories · arXiv:2402.05127
-
Is Mamba Capable of In-Context Learning? 5 Feb 2024 · 1 repository · arXiv:2402.03170Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Just Cluster It: An Approach for Exploration in High-Dimensions using Clustering and Pre-Trained Representations 5 Feb 2024 · 1 repository · arXiv:2402.03138Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
LB-KBQA: Large-language-model and BERT based Knowledge-Based Question and Answering System 5 Feb 2024 · 0 repositories · arXiv:2402.05130
-
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.02896
-
MobilityGPT: Enhanced Human Mobility Modeling with a GPT model 5 Feb 2024 · 0 repositories · arXiv:2402.03264
-
Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations 5 Feb 2024 · 0 repositories · arXiv:2402.03053
-
SWAG: Storytelling With Action Guidance 5 Feb 2024 · 1 repository · arXiv:2402.03483
-
Time-, Memory- and Parameter-Efficient Visual Adaptation 5 Feb 2024 · 0 repositories · arXiv:2402.02887Syntology 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Toward Human-AI Alignment in Large-Scale Multi-Player Games 5 Feb 2024 · 0 repositories · arXiv:2402.03575
-
UniMem: Towards a Unified View of Long-Context Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03009
-
A flexible Bayesian g-formula for causal survival analyses with time-dependent confounding 4 Feb 2024 · 1 repository · arXiv:2402.02306
-
A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer 4 Feb 2024 · 1 repository · arXiv:2402.02464Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Aligner: Efficient Alignment by Learning to Correct 4 Feb 2024 · 0 repositories · arXiv:2402.02416
-
AutoTimes: Autoregressive Time Series Forecasters via Large Language Models 4 Feb 2024 · 1 repository · arXiv:2402.02370
-
Breaking MLPerf Training: A Case Study on Optimizing BERT 4 Feb 2024 · 0 repositories · arXiv:2402.02447
-
Evaluating Large Language Models in Analysing Classroom Dialogue 4 Feb 2024 · 0 repositories · arXiv:2402.02380
-
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering 4 Feb 2024 · 1 repository · arXiv:2402.02503Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation 4 Feb 2024 · 0 repositories · arXiv:2402.14594
-
INViT: A Generalizable Routing Problem Solver with Invariant Nested View Transformer 4 Feb 2024 · 1 repository · arXiv:2402.02317Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 12 pointer-only (licence)