Methods › General › Attention Modules › Multi-Head Attention › Papers, page 34
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 34 of 249: papers 3,301 to 3,400 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding 21 Nov 2024 · 1 repository · arXiv:2411.14347
-
Evaluating the Robustness of Analogical Reasoning in Large Language Models 21 Nov 2024 · 1 repository · arXiv:2411.14215Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Explaining GPT-4's Schema of Depression Using Machine Behavior Analysis 21 Nov 2024 · 0 repositories · arXiv:2411.13800
-
FastRAG: Retrieval Augmented Generation for Semi-structured Data 21 Nov 2024 · 0 repositories · arXiv:2411.13773
-
G-RAG: Knowledge Expansion in Material Science 21 Nov 2024 · 1 repository · arXiv:2411.14592
-
Generative Fuzzy System for Sequence Generation 21 Nov 2024 · 0 repositories · arXiv:2411.13867
-
Global and Local Attention-Based Transformer for Hyperspectral Image Change Detection 21 Nov 2024 · 1 repository · arXiv:2411.14109
-
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI 21 Nov 2024 · 1 repository · arXiv:2411.14522Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly 21 Nov 2024 · 0 repositories · arXiv:2411.14121
-
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation 21 Nov 2024 · 0 repositories · arXiv:2411.15224
-
POS-tagging to highlight the skeletal structure of sentences 21 Nov 2024 · 2 repositories · arXiv:2411.14393
-
Stable Flow: Vital Layers for Training-Free Image Editing 21 Nov 2024 · 1 repository · arXiv:2411.14430Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
The Master-Slave Encoder Model for Improving Patent Text Summarization: A New Approach to Combining Specifications and Claims 21 Nov 2024 · 0 repositories · arXiv:2411.14072
-
Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective 21 Nov 2024 · 0 repositories · arXiv:2411.14572
-
Understanding World or Predicting Future? A Comprehensive Survey of World Models 21 Nov 2024 · 0 repositories · arXiv:2411.14499
-
AI-Driven Agents with Prompts Designed for High Agreeableness Increase the Likelihood of Being Mistaken for a Human in the Turing Test 20 Nov 2024 · 0 repositories · arXiv:2411.13749
-
BIPro: Zero-shot Chinese Poem Generation via Block Inverse Prompting Constrained Generation Framework 20 Nov 2024 · 0 repositories · arXiv:2411.13237
-
Combining Autoregressive and Autoencoder Language Models for Text Classification 20 Nov 2024 · 1 repository · arXiv:2411.13282
-
DMQR-RAG: Diverse Multi-Query Rewriting for RAG 20 Nov 2024 · 0 repositories · arXiv:2411.13154
-
DrugGen: Advancing Drug Discovery with Large Language Models and Reinforcement Learning Feedback 20 Nov 2024 · 4 repositories · arXiv:2411.14157
-
Exploring Large Language Models for Climate Forecasting 20 Nov 2024 · 0 repositories · arXiv:2411.13724
-
MemoryFormer: Minimize Transformer Computation by Removing Fully-Connected Layers 20 Nov 2024 · 0 repositories · arXiv:2411.12992Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Multimodal large language model for wheat breeding: a new exploration of smart breeding 20 Nov 2024 · 0 repositories · arXiv:2411.15203
-
Multipath Mitigation Technology-integrated GNSS Direct Position Estimation Plug-in Module 20 Nov 2024 · 0 repositories · arXiv:2411.13339
-
On the Way to LLM Personalization: Learning to Remember User Conversations 20 Nov 2024 · 0 repositories · arXiv:2411.13405
-
Quantum Attention for Vision Transformers in High Energy Physics 20 Nov 2024 · 0 repositories · arXiv:2411.13520
-
Retrieval-Augmented Generation for Domain-Specific Question Answering: A Case Study on Pittsburgh and CMU 20 Nov 2024 · 0 repositories · arXiv:2411.13691
-
Scaling Laws for Online Advertisement Retrieval 20 Nov 2024 · 0 repositories · arXiv:2411.13322
-
The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz 20 Nov 2024 · 0 repositories · arXiv:2411.14486
-
Transformers with Sparse Attention for Granger Causality 20 Nov 2024 · 0 repositories · arXiv:2411.13264
-
Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding 20 Nov 2024 · 0 repositories · arXiv:2411.13163
-
A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation 19 Nov 2024 · 0 repositories · arXiv:2411.12157
-
Comparing Prior and Learned Time Representations in Transformer Models of Timeseries 19 Nov 2024 · 0 repositories · arXiv:2411.12476
-
DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models 19 Nov 2024 · 1 repository · arXiv:2411.12643
-
Enhancing Multi-Class Disease Classification: Neoplasms, Cardiovascular, Nervous System, and Digestive Disorders Using Advanced LLMs 19 Nov 2024 · 0 repositories · arXiv:2411.12712
-
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages 19 Nov 2024 · 0 repositories · arXiv:2411.12240
-
Faster Multi-GPU Training with PPLL: A Pipeline Parallelism Framework Leveraging Local Learning 19 Nov 2024 · 0 repositories · arXiv:2411.12780
-
Leveraging Virtual Reality and AI Tutoring for Language Learning: A Case Study of a Virtual Campus Environment with OpenAI GPT Integration with Unity 3D 19 Nov 2024 · 0 repositories · arXiv:2411.12619
-
Multi-Grained Preference Enhanced Transformer for Multi-Behavior Sequential Recommendation 19 Nov 2024 · 1 repository · arXiv:2411.12179
-
PoM: Efficient Image and Video Generation with the Polynomial Mixer 19 Nov 2024 · 1 repository · arXiv:2411.12663
-
Residual Vision Transformer (ResViT) Based Self-Supervised Learning Model for Brain Tumor Classification 19 Nov 2024 · 0 repositories · arXiv:2411.12874
-
Strengthening Fake News Detection: Leveraging SVM and Sophisticated Text Vectorization Techniques. Defying BERT? 19 Nov 2024 · 0 repositories · arXiv:2411.12703
-
Transformer Neural Processes -- Kernel Regression 19 Nov 2024 · 0 repositories · arXiv:2411.12502
-
Ultra-Sparse Memory Network 19 Nov 2024 · 0 repositories · arXiv:2411.12364
-
Advacheck at GenAI Detection Task 1: AI Detection Powered by Domain-Aware Multi-Tasking 18 Nov 2024 · 1 repository · arXiv:2411.11736
-
Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification 18 Nov 2024 · 0 repositories · arXiv:2411.14474
-
Can Open-source LLMs Enhance Data Synthesis for Toxic Detection?: An Experimental Study 18 Nov 2024 · 0 repositories · arXiv:2411.15175
-
Chapter 7 Review of Data-Driven Generative AI Models for Knowledge Extraction from Scientific Literature in Healthcare 18 Nov 2024 · 0 repositories · arXiv:2411.11635
-
CNMBERT: A Model for Converting Hanyu Pinyin Abbreviations to Chinese Characters 18 Nov 2024 · 1 repository · arXiv:2411.11770
-
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery 18 Nov 2024 · 0 repositories · arXiv:2411.11214
-
Enhancing Decision Transformer with Diffusion-Based Trajectory Branch Generation 18 Nov 2024 · 0 repositories · arXiv:2411.11327
-
FCC: Fully Connected Correlation for Few-Shot Segmentation 18 Nov 2024 · 0 repositories · arXiv:2411.11917
-
In-Situ Melt Pool Characterization via Thermal Imaging for Defect Detection in Directed Energy Deposition Using Vision Transformers 18 Nov 2024 · 0 repositories · arXiv:2411.12028
-
LaVin-DiT: Large Vision Diffusion Transformer 18 Nov 2024 · 0 repositories · arXiv:2411.11505
-
LiTformer: Efficient Modeling and Analysis of High-Speed Link Transmitters Using Non-Autoregressive Transformer 18 Nov 2024 · 0 repositories · arXiv:2411.11699
-
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback 18 Nov 2024 · 1 repository · arXiv:2412.03578Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Popular LLMs Amplify Race and Gender Disparities in Human Mobility 18 Nov 2024 · 0 repositories · arXiv:2411.14469
-
SeqProFT: Applying LoRA Finetuning for Sequence-only Protein Property Predictions 18 Nov 2024 · 0 repositories · arXiv:2411.11530
-
ST-Tree with Interpretability for Multivariate Time Series Classification 18 Nov 2024 · 0 repositories · arXiv:2411.11620
-
Suicide Risk Assessment on Social Media with Semi-Supervised Learning 18 Nov 2024 · 0 repositories · arXiv:2411.12767
-
TimeFormer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction 18 Nov 2024 · 1 repository · arXiv:2411.11941
-
Understanding Student Sentiment on Mental Health Support in Colleges Using Large Language Models 18 Nov 2024 · 0 repositories · arXiv:2412.04326
-
Unveiling the Inflexibility of Adaptive Embedding in Traffic Forecasting 18 Nov 2024 · 1 repository · arXiv:2411.11448
-
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs 18 Nov 2024 · 1 repository · arXiv:2411.11266
-
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection 17 Nov 2024 · 1 repository · arXiv:2411.10922
-
Freqformer: Frequency-Domain Transformer for 3-D Visualization and Quantification of Human Retinal Circulation 17 Nov 2024 · 0 repositories · arXiv:2411.11189
-
IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers 17 Nov 2024 · 0 repositories · arXiv:2411.10956
-
Knowledge-enhanced Transformer for Multivariate Long Sequence Time-series Forecasting 17 Nov 2024 · 0 repositories · arXiv:2411.11046
-
A Novel Approach to Eliminating Hallucinations in Large Language Model-Assisted Causal Discovery 16 Nov 2024 · 0 repositories · arXiv:2411.12759
-
A Wearable Gait Monitoring System for 17 Gait Parameters Based on Computer Vision 16 Nov 2024 · 0 repositories · arXiv:2411.10739
-
AllRestorer: All-in-One Transformer for Image Restoration under Composite Degradations 16 Nov 2024 · 0 repositories · arXiv:2411.10708
-
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer 16 Nov 2024 · 1 repository · arXiv:2411.10781
-
IntentGPT: Few-shot Intent Discovery with Large Language Models 16 Nov 2024 · 0 repositories · arXiv:2411.10670
-
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 16 Nov 2024 · 1 repository · arXiv:2411.10741Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection 16 Nov 2024 · 1 repository · arXiv:2411.10888
-
A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission 15 Nov 2024 · 0 repositories · arXiv:2411.09936
-
Building 6G Radio Foundation Models with Transformer Architectures 15 Nov 2024 · 0 repositories · arXiv:2411.09996
-
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation 15 Nov 2024 · 0 repositories · arXiv:2411.10060
-
Debias-CLR: A Contrastive Learning Based Debiasing Method for Algorithmic Fairness in Healthcare Applications 15 Nov 2024 · 0 repositories · arXiv:2411.10544
-
DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization 15 Nov 2024 · 0 repositories · arXiv:2411.10193
-
Does Prompt Formatting Have Any Impact on LLM Performance? 15 Nov 2024 · 0 repositories · arXiv:2411.10541
-
Evidential Federated Learning for Skin Lesion Image Classification 15 Nov 2024 · 0 repositories · arXiv:2411.10071
-
Hysteresis Activation Function for Efficient Inference 15 Nov 2024 · 1 repository · arXiv:2411.10573Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models? 15 Nov 2024 · 1 repository · arXiv:2411.10020
-
LoRA-LiteE: A Computationally Efficient Framework for Chatbot Preference-Tuning 15 Nov 2024 · 0 repositories · arXiv:2411.09947
-
MARS: Unleashing the Power of Variance Reduction for Training Large Models 15 Nov 2024 · 2 repositories · arXiv:2411.10438Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Probabilistic Prior Driven Attention Mechanism Based on Diffusion Model for Imaging Through Atmospheric Turbulence 15 Nov 2024 · 0 repositories · arXiv:2411.10321
-
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation 15 Nov 2024 · 0 repositories · arXiv:2411.10129
-
RETR: Multi-View Radar Detection Transformer for Indoor Perception 15 Nov 2024 · 1 repository · arXiv:2411.10293Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
P² Law: Scaling Law for Post-Training After Model Pruning 15 Nov 2024 · 0 repositories · arXiv:2411.10272
-
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism 15 Nov 2024 · 0 repositories · arXiv:2411.10543
-
Take Package as Language: Anomaly Detection Using Transformer 15 Nov 2024 · 0 repositories · arXiv:2412.04473
-
ULTra: Unveiling Latent Token Interpretability in Transformer Based Understanding 15 Nov 2024 · 0 repositories · arXiv:2411.12589
-
Adopting RAG for LLM-Aided Future Vehicle Design 14 Nov 2024 · 0 repositories · arXiv:2411.09590
-
Assessing the Performance of the DINOv2 Self-supervised Learning Vision Transformer Model for the Segmentation of the Left Atrium from MRI Images 14 Nov 2024 · 0 repositories · arXiv:2411.09598
-
Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming 14 Nov 2024 · 0 repositories · arXiv:2411.09261
-
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency 14 Nov 2024 · 0 repositories · arXiv:2411.09587
-
Beyond Static Tools: Evaluating Large Language Models for Cryptographic Misuse Detection 14 Nov 2024 · 0 repositories · arXiv:2411.09772
-
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering 14 Nov 2024 · 0 repositories · arXiv:2411.09213
-
DSCformer: A Dual-Branch Network Integrating Enhanced Dynamic Snake Convolution and SegFormer for Crack Segmentation 14 Nov 2024 · 0 repositories · arXiv:2411.09371