Methods › General › Attention Modules › Multi-Head Attention › Papers, page 45
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 45 of 249: papers 4,401 to 4,500 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
RATIONALYST: Pre-training Process-Supervision for Improving Reasoning 1 Oct 2024 · 1 repository · arXiv:2410.01044
-
Robust Traffic Forecasting against Spatial Shift over Years 1 Oct 2024 · 1 repository · arXiv:2410.00373Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Sparse Attention Decomposition Applied to Circuit Tracing 1 Oct 2024 · 1 repository · arXiv:2410.00340
-
STGformer: Efficient Spatiotemporal Graph Transformer for Traffic Forecasting 1 Oct 2024 · 1 repository · arXiv:2410.00385
-
TFCT-I2P: Three stream fusion network with color aware transformer for image-to-point cloud registration 1 Oct 2024 · 1 repository · arXiv:2410.00360
-
TransResNet: Integrating the Strengths of ViTs and CNNs for High Resolution Medical Image Segmentation via Feature Grafting 1 Oct 2024 · 1 repository · arXiv:2410.00986
-
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions 30 Sep 2024 · 0 repositories · arXiv:2409.20303
-
A Methodology for Explainable Large Language Models with Integrated Gradients and Linguistic Analysis in Text Classification 30 Sep 2024 · 0 repositories · arXiv:2410.00250
-
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer 30 Sep 2024 · 0 repositories · arXiv:2410.00086
-
Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation 30 Sep 2024 · 0 repositories · arXiv:2410.00163
-
ASQuery: A Query-based Model for Action Segmentation 30 Sep 2024 · 1 repository
-
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain 30 Sep 2024 · 0 repositories · arXiv:2409.20075
-
CBAM-SwinT-BL: Small Rail Surface Defect Detection Method Based on Swin Transformer with Block Level CBAM Enhancement 30 Sep 2024 · 0 repositories · arXiv:2409.20113
-
CliMB: An AI-enabled Partner for Clinical Predictive Modeling 30 Sep 2024 · 1 repository · arXiv:2410.03736Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Depression detection in social media posts using transformer-based models and auxiliary features 30 Sep 2024 · 0 repositories · arXiv:2409.20048
-
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification 30 Sep 2024 · 1 repository · arXiv:2410.00179
-
GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction 30 Sep 2024 · 0 repositories · arXiv:2409.20223
-
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG 30 Sep 2024 · 0 repositories · arXiv:2410.02825
-
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation 30 Sep 2024 · 0 repositories · arXiv:2409.19937
-
Modelando procesos cognitivos de la lectura natural con GPT-2 30 Sep 2024 · 0 repositories · arXiv:2409.20174
-
Exploring Social Media Image Categorization Using Large Models with Different Adaptation Methods: A Case Study on Cultural Nature's Contributions to People 30 Sep 2024 · 0 repositories · arXiv:2410.00275
-
On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability 30 Sep 2024 · 2 repositories · arXiv:2409.19924
-
QAEncoder: Towards Aligned Representation Learning in Question Answering System 30 Sep 2024 · 1 repository · arXiv:2409.20434
-
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels 30 Sep 2024 · 0 repositories · arXiv:2409.19846
-
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers 29 Sep 2024 · 0 repositories · arXiv:2409.19566
-
Adversarial Examples for DNA Classification 29 Sep 2024 · 0 repositories · arXiv:2409.19788
-
Can Models Learn Skill Composition from Examples? 29 Sep 2024 · 0 repositories · arXiv:2409.19808
-
Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional Noises 29 Sep 2024 · 0 repositories · arXiv:2409.19619
-
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems 29 Sep 2024 · 1 repository · arXiv:2409.19804Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks 29 Sep 2024 · 0 repositories · arXiv:2409.19521
-
DATransNet: Dynamic Attention Transformer Network for Infrared Small Target Detection 29 Sep 2024 · 1 repository · arXiv:2409.19599
-
InfantCryNet: A Data-driven Framework for Intelligent Analysis of Infant Cries 29 Sep 2024 · 0 repositories · arXiv:2409.19689
-
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models 29 Sep 2024 · 0 repositories · arXiv:2409.19492
-
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead 29 Sep 2024 · 0 repositories · arXiv:2409.19745
-
See then Tell: Enhancing Key Information Extraction with Vision Grounding 29 Sep 2024 · 0 repositories · arXiv:2409.19573
-
Spiking Transformer with Spatial-Temporal Attention 29 Sep 2024 · 1 repository · arXiv:2409.19764
-
Analog In-Memory Computing Attention Mechanism for Fast and Energy-Efficient Large Language Models 28 Sep 2024 · 1 repository · arXiv:2409.19315
-
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning 28 Sep 2024 · 0 repositories · arXiv:2409.19255
-
Efficient Federated Intrusion Detection in 5G ecosystem using optimized BERT-based model 28 Sep 2024 · 1 repository · arXiv:2409.19390
-
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Large Language Models and Ensemble Learning 28 Sep 2024 · 2 repositories · arXiv:2409.19467
-
Multi-Atlas Brain Network Classification through Consistency Distillation and Complementary Information Fusion 28 Sep 2024 · 0 repositories · arXiv:2410.08228
-
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization 28 Sep 2024 · 0 repositories · arXiv:2409.19345
-
AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow 27 Sep 2024 · 0 repositories · arXiv:2409.18924
-
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations 27 Sep 2024 · 0 repositories · arXiv:2409.18764
-
Cottention: Linear Transformers With Cosine Attention 27 Sep 2024 · 1 repository · arXiv:2409.18747Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture 27 Sep 2024 · 0 repositories · arXiv:2409.18568
-
How Effective is Pre-training of Large Masked Autoencoders for Downstream Earth Observation Tasks? 27 Sep 2024 · 0 repositories · arXiv:2409.18536
-
Improving Visual Object Tracking through Visual Prompting 27 Sep 2024 · 1 repository · arXiv:2409.18901
-
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction 27 Sep 2024 · 1 repository · arXiv:2409.18957
-
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning 27 Sep 2024 · 0 repositories · arXiv:2409.19075
-
Multi-Source Hard and Soft Information Fusion Approach for Accurate Cryptocurrency Price Movement Prediction 27 Sep 2024 · 0 repositories · arXiv:2409.18895
-
Not the Silver Bullet: LLM-enhanced Programming Error Messages are Ineffective in Practice 27 Sep 2024 · 0 repositories · arXiv:2409.18661
-
On the Power of Decision Trees in Auto-Regressive Language Modeling 27 Sep 2024 · 0 repositories · arXiv:2409.19150
-
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs 27 Sep 2024 · 0 repositories · arXiv:2409.18794
-
Pruning then Reweighting: Towards Data-Efficient Training of Diffusion Models 27 Sep 2024 · 1 repository · arXiv:2409.19128
-
Query matching for spatio-temporal action detection with query-based object detector 27 Sep 2024 · 0 repositories · arXiv:2409.18408
-
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models 27 Sep 2024 · 0 repositories · arXiv:2409.18654
-
Suicide Phenotyping from Clinical Notes in Safety-Net Psychiatric Hospital Using Multi-Label Classification with Pre-Trained Language Models 27 Sep 2024 · 0 repositories · arXiv:2409.18878
-
A Fuzzy-based Approach to Predict Human Interaction by Functional Near-Infrared Spectroscopy 26 Sep 2024 · 0 repositories · arXiv:2409.17661
-
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing 26 Sep 2024 · 1 repository · arXiv:2409.17453
-
CASPFormer: Trajectory Prediction from BEV Images with Deformable Attention 26 Sep 2024 · 0 repositories · arXiv:2409.17790
-
Comparing Unidirectional, Bidirectional, and Word2vec Models for Discovering Vulnerabilities in Compiled Lifted Code 26 Sep 2024 · 0 repositories · arXiv:2409.17513
-
DARE: Diverse Visual Question Answering with Robustness Evaluation 26 Sep 2024 · 0 repositories · arXiv:2409.18023
-
Developing a Dual-Stage Vision Transformer Model for Lung Disease Classification 26 Sep 2024 · 0 repositories · arXiv:2409.18257
-
Dynamic Subframe Splitting and Spatio-Temporal Motion Entangled Sparse Attention for RGB-E Tracking 26 Sep 2024 · 0 repositories · arXiv:2409.17560
-
Efficient In-Domain Question Answering for Resource-Constrained Environments 26 Sep 2024 · 0 repositories · arXiv:2409.17648
-
EM-Net: Efficient Channel and Frequency Learning with Mamba for 3D Medical Image Segmentation 26 Sep 2024 · 1 repository · arXiv:2409.17675Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation 26 Sep 2024 · 0 repositories · arXiv:2409.18313
-
Enhancing Tourism Recommender Systems for Sustainable City Trips Using Retrieval-Augmented Generation 26 Sep 2024 · 0 repositories · arXiv:2409.18003
-
HydraViT: Stacking Heads for a Scalable ViT 26 Sep 2024 · 1 repository · arXiv:2409.17978Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 16 harvested samples)
-
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization 26 Sep 2024 · 0 repositories · arXiv:2409.17534
-
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models 26 Sep 2024 · 1 repository · arXiv:2409.17481Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
MultiClimate: Multimodal Stance Detection on Climate Change Videos 26 Sep 2024 · 1 repository · arXiv:2409.18346Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
NeuroPath: A Neural Pathway Transformer for Joining the Dots of Human Connectomes 26 Sep 2024 · 1 repository · arXiv:2409.17510Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Ophthalmic Biomarker Detection with Parallel Prediction of Transformer and Convolutional Architecture 26 Sep 2024 · 0 repositories · arXiv:2409.17788
-
PEDRO: Parameter-Efficient Fine-tuning with Prompt DEpenDent Representation MOdification 26 Sep 2024 · 0 repositories · arXiv:2409.17834
-
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods 26 Sep 2024 · 0 repositories · arXiv:2409.17939
-
Retrospective Comparative Analysis of Prostate Cancer In-Basket Messages: Responses from Closed-Domain LLM vs. Clinical Teams 26 Sep 2024 · 1 repository · arXiv:2409.18290
-
Self-supervised Monocular Depth Estimation with Large Kernel Attention 26 Sep 2024 · 0 repositories · arXiv:2409.17895
-
Self-supervised Pretraining for Cardiovascular Magnetic Resonance Cine Segmentation 26 Sep 2024 · 1 repository · arXiv:2409.18100
-
T3: A Novel Zero-shot Transfer Learning Framework Iteratively Training on an Assistant Task for a Target Task 26 Sep 2024 · 0 repositories · arXiv:2409.17640
-
The application of GPT-4 in grading design university students' assignment and providing feedback: An exploratory study 26 Sep 2024 · 0 repositories · arXiv:2409.17698
-
Unifying Dimensions: A Linear Adaptive Approach to Lightweight Image Super-Resolution 26 Sep 2024 · 1 repository · arXiv:2409.17597
-
A Prompting-Based Representation Learning Method for Recommendation with Large Language Models 25 Sep 2024 · 0 repositories · arXiv:2409.16674
-
Beyond Turing Test: Can GPT-4 Sway Experts' Decisions? 25 Sep 2024 · 0 repositories · arXiv:2409.16710
-
Block Expanded DINORET: Adapting Natural Domain Foundation Models for Retinal Imaging Without Catastrophic Forgetting 25 Sep 2024 · 0 repositories · arXiv:2409.17332
-
CodeInsight: A Curated Dataset of Practical Coding Solutions from Stack Overflow 25 Sep 2024 · 1 repository · arXiv:2409.16819
-
Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Handy Appetizer 25 Sep 2024 · 0 repositories · arXiv:2409.17120
-
Enhancing Automatic Keyphrase Labelling with Text-to-Text Transfer Transformer (T5) Architecture: A Framework for Keyphrase Generation and Filtering 25 Sep 2024 · 0 repositories · arXiv:2409.16760
-
Going Beyond U-Net: Assessing Vision Transformers for Semantic Segmentation in Microscopy Image Analysis 25 Sep 2024 · 0 repositories · arXiv:2409.16940
-
Gradient Boosting Decision Trees on Medical Diagnosis over Tabular Data 25 Sep 2024 · 1 repository · arXiv:2410.03705
-
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space 25 Sep 2024 · 1 repository · arXiv:2409.16897
-
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents 25 Sep 2024 · 1 repository · arXiv:2409.16934
-
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ 25 Sep 2024 · 0 repositories · arXiv:2409.16779
-
Non-stationary BERT: Exploring Augmented IMU Data For Robust Human Activity Recognition 25 Sep 2024 · 0 repositories · arXiv:2409.16730
-
Post-hoc Reward Calibration: A Case Study on Length Bias 25 Sep 2024 · 1 repository · arXiv:2409.17407Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Pre-trained Graphformer-based Ranking at Web-scale Search (Extended Abstract) 25 Sep 2024 · 0 repositories · arXiv:2409.16590
-
Probing Omissions and Distortions in Transformer-based RDF-to-Text Models 25 Sep 2024 · 0 repositories · arXiv:2409.16707
-
Quantum-Classical Sentiment Analysis 25 Sep 2024 · 0 repositories · arXiv:2409.16928
-
Severity Prediction in Mental Health: LLM-based Creation, Analysis, Evaluation of a Novel Multilingual Dataset 25 Sep 2024 · 0 repositories · arXiv:2409.17397