Methods › General › Attention Modules › Multi-Head Attention › Papers, page 104
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 104 of 249: papers 10,301 to 10,400 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Drilling Down into the Discourse Structure with LLMs for Long Document Question Answering 22 Nov 2023 · 0 repositories · arXiv:2311.13565
-
Generation of Explanations for Logic Reasoning 22 Nov 2023 · 0 repositories · arXiv:2311.13455
-
HEViTPose: High-Efficiency Vision Transformer for Human Pose Estimation 22 Nov 2023 · 1 repository · arXiv:2311.13615
-
Input Compression with Positional Consistency for Efficient Training and Inference of Transformer Neural Networks 22 Nov 2023 · 1 repository · arXiv:2312.12385
-
Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning 22 Nov 2023 · 0 repositories · arXiv:2311.13721
-
PG-Video-LLaVA: Pixel Grounding Large Video-Language Models 22 Nov 2023 · 1 repository · arXiv:2311.13435Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation 22 Nov 2023 · 1 repository · arXiv:2311.13602
-
AlignedCoT: Prompting Large Language Models via Native-Speaking Demonstrations 22 Nov 2023 · 1 repository · arXiv:2311.13538
-
Surpassing GPT-4 Medical Coding with a Two-Stage Approach 22 Nov 2023 · 0 repositories · arXiv:2311.13735
-
Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs 22 Nov 2023 · 1 repository · arXiv:2311.13194Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Unified Classification and Rejection: A One-versus-All Framework 22 Nov 2023 · 1 repository · arXiv:2311.13355Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
@ve: A Chatbot for Latin 22 Nov 2023 · 0 repositories · arXiv:2311.14741
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? 22 Nov 2023 · 1 repository · arXiv:2311.13110Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Survey on Large Language Models for Personalized and Explainable Recommendations 21 Nov 2023 · 0 repositories · arXiv:2311.12338
-
AcademicGPT: Empowering Academic Research 21 Nov 2023 · 0 repositories · arXiv:2311.12315
-
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey 21 Nov 2023 · 1 repository · arXiv:2311.12351
-
ALPHA: AnomaLous Physiological Health Assessment Using Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.12524
-
GeoLocator: a location-integrated large multimodal model for inferring geo-privacy 21 Nov 2023 · 0 repositories · arXiv:2311.13018
-
AudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning 21 Nov 2023 · 1 repository · arXiv:2311.12371
-
Causality is all you need 21 Nov 2023 · 0 repositories · arXiv:2311.12307
-
Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning 21 Nov 2023 · 1 repository · arXiv:2311.13612
-
Extracting Definienda in Mathematical Scholarly Articles with Transformers 21 Nov 2023 · 2 repositories · arXiv:2311.12448
-
From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.13063
-
GAIA: a benchmark for General AI Assistants 21 Nov 2023 · 2 repositories · arXiv:2311.12983
-
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning 21 Nov 2023 · 0 repositories · arXiv:2311.12631
-
HoVer-UNet: Accelerating HoVerNet with UNet-based multi-class nuclei segmentation via knowledge distillation 21 Nov 2023 · 1 repository · arXiv:2311.12553
-
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks 21 Nov 2023 · 1 repository · arXiv:2311.12997Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
HPCNeuroNet: Advancing Neuromorphic Audio Signal Processing with Transformer-Enhanced Spiking Neural Networks 21 Nov 2023 · 0 repositories · arXiv:2311.12449
-
IEKM: A Model Incorporating External Keyword Matrices 21 Nov 2023 · 0 repositories · arXiv:2311.12310
-
Improving Source-Free Target Adaptation with Vision Transformers Leveraging Domain Representation Images 21 Nov 2023 · 0 repositories · arXiv:2311.12589
-
Interpretation of the Transformer and Improvement of the Extractor 21 Nov 2023 · 1 repository · arXiv:2311.12678
-
InterPrompt: Interpretable Prompting for Interrelated Interpersonal Risk Factors in Reddit Posts 21 Nov 2023 · 0 repositories · arXiv:2311.12404
-
Learning to Compute Gröbner Bases 21 Nov 2023 · 2 repositories · arXiv:2311.12904Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 17 pointer-only (licence)
-
Learning to Optimise Wind Farms with Graph Transformers 21 Nov 2023 · 0 repositories · arXiv:2311.12750
-
Long-MIL: Scaling Long Contextual Multiple Instance Learning for Histopathology Whole Slide Image Analysis 21 Nov 2023 · 0 repositories · arXiv:2311.12885
-
LowResource at BLP-2023 Task 2: Leveraging BanglaBert for Low Resource Sentiment Analysis of Bangla Language 21 Nov 2023 · 1 repository · arXiv:2311.12735
-
Oasis: Data Curation and Assessment System for Pretraining of Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.12537
-
Residual Aligner-based Network (RAN): Motion-separable structure for coarse-to-fine discontinuous deformable registration 21 Nov 2023 · 1 repository
-
Utilizing Language Models for Tour Itinerary Recommendation 21 Nov 2023 · 0 repositories · arXiv:2311.12355
-
A Large-Scale Car Parts (LSCP) Dataset for Lightweight Fine-Grained Detection 20 Nov 2023 · 0 repositories · arXiv:2311.11754
-
A novel transformer-based approach for soil temperature prediction 20 Nov 2023 · 0 repositories · arXiv:2311.11626
-
Assessing Prompt Injection Risks in 200+ Custom GPTs 20 Nov 2023 · 1 repository · arXiv:2311.11538
-
Correlated Attention in Transformers for Multivariate Time Series 20 Nov 2023 · 0 repositories · arXiv:2311.11959
-
Decoupled DETR For Few-shot Object Detection 20 Nov 2023 · 0 repositories · arXiv:2311.11570
-
Disentangling Structure and Appearance in ViT Feature Space 20 Nov 2023 · 0 repositories · arXiv:2311.12193
-
Evil Geniuses: Delving into the Safety of LLM-based Agents 20 Nov 2023 · 1 repository · arXiv:2311.11855Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
FreeKD: Knowledge Distillation via Semantic Frequency Prompt 20 Nov 2023 · 1 repository · arXiv:2311.12079Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Generating Valid and Natural Adversarial Examples with Large Language Models 20 Nov 2023 · 0 repositories · arXiv:2311.11861
-
GPQA: A Graduate-Level Google-Proof Q&A Benchmark 20 Nov 2023 · 3 repositories · arXiv:2311.12022Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents 20 Nov 2023 · 1 repository · arXiv:2311.11844
-
LiDAR-HMR: 3D Human Mesh Recovery from LiDAR 20 Nov 2023 · 2 repositories · arXiv:2311.11971
-
LogLead -- Fast and Integrated Log Loader, Enhancer, and Anomaly Detector 20 Nov 2023 · 1 repository · arXiv:2311.11809
-
LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning 20 Nov 2023 · 1 repository · arXiv:2311.12023Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
MemoryCompanion: A Smart Healthcare Solution to Empower Efficient Alzheimer's Care Via Unleashing Generative AI 20 Nov 2023 · 0 repositories · arXiv:2311.14730
-
Meta Prompting for AI Systems 20 Nov 2023 · 1 repository · arXiv:2311.11482Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MGCT: Mutual-Guided Cross-Modality Transformer for Survival Outcome Prediction using Integrative Histopathology-Genomic Features 20 Nov 2023 · 1 repository · arXiv:2311.11659
-
PMP-Swin: Multi-Scale Patch Message Passing Swin Transformer for Retinal Disease Classification 20 Nov 2023 · 0 repositories · arXiv:2311.11669
-
Refactoring Programs Using Large Language Models with Few-Shot Examples 20 Nov 2023 · 0 repositories · arXiv:2311.11690
-
Tiny-VBF: Resource-Efficient Vision Transformer based Lightweight Beamformer for Ultrasound Single-Angle Plane Wave Imaging 20 Nov 2023 · 0 repositories · arXiv:2311.12082
-
Unveiling the Power of Self-Attention for Shipping Cost Prediction: The Rate Card Transformer 20 Nov 2023 · 1 repository · arXiv:2311.11694
-
Which AI Technique Is Better to Classify Requirements? An Experiment with SVM, LSTM, and ChatGPT 20 Nov 2023 · 1 repository · arXiv:2311.11547
-
Inspecting Explainability of Transformer Models with Additional Statistical Information 19 Nov 2023 · 0 repositories · arXiv:2311.11378
-
Shape-Sensitive Loss for Catheter and Guidewire Segmentation 19 Nov 2023 · 0 repositories · arXiv:2311.11205
-
Spot the Bot: Distinguishing Human-Written and Bot-Generated Texts Using Clustering and Information Theory Techniques 19 Nov 2023 · 0 repositories · arXiv:2311.11441
-
Tensor-Aware Energy Accounting 19 Nov 2023 · 1 repository · arXiv:2311.11424
-
Behavior Optimized Image Generation 18 Nov 2023 · 0 repositories · arXiv:2311.10995
-
Bit Cipher -- A Simple yet Powerful Word Representation System that Integrates Efficiently with Language Models 18 Nov 2023 · 0 repositories · arXiv:2311.11012
-
Compositional Fusion of Signals in Data Embedding 18 Nov 2023 · 0 repositories · arXiv:2311.11085
-
Partially Randomizing Transformer Weights for Dialogue Response Diversity 18 Nov 2023 · 0 repositories · arXiv:2311.10943
-
Structure-Aware Sparse-View X-ray 3D Reconstruction 18 Nov 2023 · 2 repositories · arXiv:2311.10959Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language 18 Nov 2023 · 1 repository · arXiv:2311.11142
-
Visual AI and Linguistic Intelligence Through Steerability and Composability 18 Nov 2023 · 0 repositories · arXiv:2312.12383
-
Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers 17 Nov 2023 · 0 repositories · arXiv:2311.10242
-
Enhancing Data Efficiency and Feature Identification for Lithium-Ion Battery Lifespan Prediction by Deciphering Interpretation of Temporal Patterns and Cyclic Variability Using Attention-Based Models 17 Nov 2023 · 0 repositories · arXiv:2311.10792
-
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads 17 Nov 2023 · 0 repositories · arXiv:2311.10395
-
Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2 17 Nov 2023 · 3 repositories · arXiv:2311.10702
-
DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines 17 Nov 2023 · 2 repositories · arXiv:2311.10418
-
EduQuick: A Dataset Toward Evaluating Summarization of Informal Educational Content for Social Media 17 Nov 2023 · 0 repositories
-
Extracting periodontitis diagnosis in clinical notes with RoBERTa and regular expression 17 Nov 2023 · 0 repositories · arXiv:2311.10809
-
Hashing it Out: Predicting Unhealthy Conversations on Twitter 17 Nov 2023 · 1 repository · arXiv:2311.10596
-
Multi-entity Video Transformers for Fine-Grained Video Representation Learning 17 Nov 2023 · 1 repository · arXiv:2311.10873
-
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers 17 Nov 2023 · 0 repositories · arXiv:2311.10642
-
Semi-supervised ViT knowledge distillation network with style transfer normalization for colorectal liver metastases survival prediction 17 Nov 2023 · 0 repositories · arXiv:2311.10305
-
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes 17 Nov 2023 · 1 repository · arXiv:2311.10797
-
Use GPT-J Prompt Generation with RoBERTa for NER Models on Diagnosis Extraction of Periodontal Diagnosis from Electronic Dental Records 17 Nov 2023 · 0 repositories · arXiv:2311.10810
-
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems 16 Nov 2023 · 1 repository · arXiv:2311.09476Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
BLT: Can Large Language Models Handle Basic Legal Text? 16 Nov 2023 · 1 repository · arXiv:2311.09693
-
Co-data Learning for Bayesian Additive Regression Trees 16 Nov 2023 · 1 repository · arXiv:2311.09997
-
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization 16 Nov 2023 · 0 repositories · arXiv:2311.09559
-
Event Causality Is Key to Computational Story Understanding 16 Nov 2023 · 1 repository · arXiv:2311.09648Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability 16 Nov 2023 · 0 repositories · arXiv:2311.09696
-
GEE! Grammar Error Explanation with Large Language Models 16 Nov 2023 · 1 repository · arXiv:2311.09517
-
Generative AI for Hate Speech Detection: Evaluation and Findings 16 Nov 2023 · 0 repositories · arXiv:2311.09993
-
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs 16 Nov 2023 · 1 repository · arXiv:2311.09774Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks 16 Nov 2023 · 0 repositories · arXiv:2311.09825
-
Improved TokenPose with Sparsity 16 Nov 2023 · 0 repositories · arXiv:2311.09653
-
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair 16 Nov 2023 · 1 repository · arXiv:2311.09868
-
Investigating Data Contamination in Modern Benchmarks for Large Language Models 16 Nov 2023 · 0 repositories · arXiv:2311.09783
-
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains 16 Nov 2023 · 1 repository · arXiv:2311.09797Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Large Language Models for Propaganda Span Annotation 16 Nov 2023 · 1 repository · arXiv:2311.09812