Methods › General › Attention Mechanisms › Attention › Papers, page 159
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 159 of 316: papers 15,801 to 15,900 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Improving Semantic Control in Discrete Latent Spaces with Transformer Quantized Variational Autoencoders 1 Feb 2024 · 1 repository · arXiv:2402.00723
-
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing 1 Feb 2024 · 0 repositories · arXiv:2402.00658
-
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts 1 Feb 2024 · 1 repository · arXiv:2402.00433Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Multivariate Probabilistic Time Series Forecasting with Correlated Errors 1 Feb 2024 · 1 repository · arXiv:2402.01000Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Ocassionally Secure: A Comparative Analysis of Code Generation Assistants 1 Feb 2024 · 0 repositories · arXiv:2402.00689
-
On the Psychology of GPT-4: Moderately anxious, slightly masculine, honest, and humble 1 Feb 2024 · 0 repositories · arXiv:2402.01777
-
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models 1 Feb 2024 · 1 repository · arXiv:2402.00794
-
Investigating Recurrent Transformers with Dynamic Halt 1 Feb 2024 · 1 repository · arXiv:2402.00976
-
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes 1 Feb 2024 · 0 repositories · arXiv:2402.00987
-
SPARQL Generation with Entity Pre-trained GPT for KG Question Answering 1 Feb 2024 · 1 repository · arXiv:2402.00969
-
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization? 1 Feb 2024 · 0 repositories · arXiv:2402.00841
-
Human-mediated Large Language Models for Robotic Intervention in Children with Autism Spectrum Disorders 1 Feb 2024 · 0 repositories · arXiv:2402.00260
-
Ultra Fast Transformers on FPGAs for Particle Physics Experiments 1 Feb 2024 · 0 repositories · arXiv:2402.01047
-
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling 1 Feb 2024 · 0 repositories · arXiv:2402.00522
-
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM 31 Jan 2024 · 0 repositories · arXiv:2402.00097
-
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition 31 Jan 2024 · 0 repositories · arXiv:2401.17604
-
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters 31 Jan 2024 · 1 repository · arXiv:2402.10930
-
Document Structure in Long Document Transformers 31 Jan 2024 · 0 repositories · arXiv:2401.17658
-
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning 31 Jan 2024 · 1 repository · arXiv:2401.17690
-
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information 31 Jan 2024 · 0 repositories · arXiv:2401.17981
-
Exploring the limits of decoder-only models trained on public speech recognition corpora 31 Jan 2024 · 0 repositories · arXiv:2402.00235
-
Global-Liar: Factuality of LLMs over Time and Geographic Regions 31 Jan 2024 · 0 repositories · arXiv:2401.17839
-
Graph Transformers without Positional Encodings 31 Jan 2024 · 0 repositories · arXiv:2401.17791
-
Head and Neck Tumor Segmentation from [18F]F-FDG PET/CT Images Based on 3D Diffusion Model 31 Jan 2024 · 0 repositories · arXiv:2401.17593
-
Leveraging Swin Transformer for Local-to-Global Weakly Supervised Semantic Segmentation 31 Jan 2024 · 1 repository · arXiv:2401.17828
-
LLM Voting: Human Choices and AI Collective Decision Making 31 Jan 2024 · 1 repository · arXiv:2402.01766Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Local Feature Matching Using Deep Learning: A Survey 31 Jan 2024 · 1 repository · arXiv:2401.17592
-
Making a Long Story Short in Conversation Modeling 31 Jan 2024 · 0 repositories · arXiv:2402.00143
-
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding 31 Jan 2024 · 0 repositories · arXiv:2401.17692
-
Paramanu: A Family of Novel Efficient Generative Foundation Language Models for Indian Languages 31 Jan 2024 · 0 repositories · arXiv:2401.18034
-
Positional Encoding Helps Recurrent Neural Networks Handle a Large Vocabulary 31 Jan 2024 · 1 repository · arXiv:2402.00236
-
RAG-Fusion: a New Take on Retrieval-Augmented Generation 31 Jan 2024 · 0 repositories · arXiv:2402.03367
-
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval 31 Jan 2024 · 3 repositories · arXiv:2401.18059Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability 31 Jan 2024 · 0 repositories · arXiv:2402.00901
-
SCAPE: Searching Conceptual Architecture Prompts using Evolution 31 Jan 2024 · 1 repository · arXiv:2402.00089
-
Scavenging Hyena: Distilling Transformers into Long Convolution Models 31 Jan 2024 · 0 repositories · arXiv:2401.17574
-
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment 31 Jan 2024 · 0 repositories · arXiv:2401.18028
-
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering 31 Jan 2024 · 1 repository · arXiv:2401.17809Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Uncertainty-Aware Explainable Recommendation with Large Language Models 31 Jan 2024 · 0 repositories · arXiv:2402.03366
-
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts 31 Jan 2024 · 1 repository · arXiv:2401.17703
-
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models 30 Jan 2024 · 0 repositories · arXiv:2401.16765
-
A Preliminary Study on Using Large Language Models in Software Pentesting 30 Jan 2024 · 0 repositories · arXiv:2401.17459
-
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter 30 Jan 2024 · 0 repositories · arXiv:2401.17373
-
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs 30 Jan 2024 · 1 repository · arXiv:2401.16638
-
CAFCT-Net: A CNN-Transformer Hybrid Network with Contextual and Attentional Feature Fusion for Liver Tumor Segmentation 30 Jan 2024 · 0 repositories · arXiv:2401.16886
-
Conditional and Modal Reasoning in Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17169
-
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17043Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Detecting mental disorder on social media: a ChatGPT-augmented explainable approach 30 Jan 2024 · 1 repository · arXiv:2401.17477
-
Detecting Racist Text in Bengali: An Ensemble Deep Learning Framework 30 Jan 2024 · 0 repositories · arXiv:2401.16748
-
Engineering A Large Language Model From Scratch 30 Jan 2024 · 0 repositories · arXiv:2401.16736
-
Fine-tuning Transformer-based Encoder for Turkish Language Understanding Tasks 30 Jan 2024 · 0 repositories · arXiv:2401.17396
-
Large Multi-Modal Models (LMMs) as Universal Foundation Models for AI-Native Wireless Systems 30 Jan 2024 · 0 repositories · arXiv:2402.01748
-
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation 30 Jan 2024 · 1 repository · arXiv:2401.17244Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Modeling how and why aquatic vegetation removal can free rural households from poverty-disease traps 30 Jan 2024 · 1 repository · arXiv:2401.17384
-
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.16745Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
OptiState: State Estimation of Legged Robots using Gated Networks with Transformer-based Vision and Kalman Filtering 30 Jan 2024 · 1 repository · arXiv:2401.16719
-
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer 30 Jan 2024 · 1 repository · arXiv:2401.16658
-
Performance Assessment of ChatGPT vs Bard in Detecting Alzheimer's Dementia 30 Jan 2024 · 0 repositories · arXiv:2402.01751
-
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks 30 Jan 2024 · 1 repository · arXiv:2401.17263Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Single Word Change is All You Need: Designing Attacks and Defenses for Text Classifiers 30 Jan 2024 · 0 repositories · arXiv:2401.17196
-
Superiority of Multi-Head Attention in In-Context Linear Regression 30 Jan 2024 · 0 repositories · arXiv:2401.17426
-
Synthetic Dialogue Dataset Generation using LLM Agents 30 Jan 2024 · 1 repository · arXiv:2401.17461
-
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives 30 Jan 2024 · 0 repositories · arXiv:2401.16677
-
Towards Generating Informative Textual Description for Neurons in Language Models 30 Jan 2024 · 0 repositories · arXiv:2401.16731
-
ViTree: Single-path Neural Tree for Step-wise Interpretable Fine-grained Visual Categorization 30 Jan 2024 · 0 repositories · arXiv:2401.17050
-
Weaver: Foundation Models for Creative Writing 30 Jan 2024 · 0 repositories · arXiv:2401.17268
-
3DG: A Framework for Using Generative AI for Handling Sparse Learner Performance Data From Intelligent Tutoring Systems 29 Jan 2024 · 1 repository · arXiv:2402.01746
-
Enhancing Topological Dependencies in Spatio-Temporal Graphs with Cycle Message Passing Blocks 29 Jan 2024 · 1 repository · arXiv:2401.15894
-
A Survey on Structure-Preserving Graph Transformers 29 Jan 2024 · 0 repositories · arXiv:2401.16176
-
Context-Former: Stitching via Latent Conditioned Sequence Modeling 29 Jan 2024 · 0 repositories · arXiv:2401.16452
-
Credit Risk Meets Large Language Models: Building a Risk Indicator from Loan Descriptions in P2P Lending 29 Jan 2024 · 0 repositories · arXiv:2401.16458
-
Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model 29 Jan 2024 · 0 repositories · arXiv:2401.16280
-
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties 29 Jan 2024 · 0 repositories · arXiv:2402.01741
-
Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report 29 Jan 2024 · 0 repositories · arXiv:2402.01733
-
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation 29 Jan 2024 · 0 repositories · arXiv:2401.16558
-
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining 29 Jan 2024 · 0 repositories · arXiv:2401.15861
-
E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models 29 Jan 2024 · 1 repository · arXiv:2401.15927Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Hybrid Transformer and Spatial-Temporal Self-Supervised Learning for Long-term Traffic Prediction 29 Jan 2024 · 0 repositories · arXiv:2401.16453
-
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports 29 Jan 2024 · 0 repositories · arXiv:2401.16578
-
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs 29 Jan 2024 · 0 repositories · arXiv:2401.16160
-
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning 29 Jan 2024 · 0 repositories · arXiv:2401.16185
-
Multi-class Regret Detection in Hindi Devanagari Script 29 Jan 2024 · 0 repositories · arXiv:2401.16561
-
Prompt4Vis: Prompting Large Language Models with Example Mining and Schema Filtering for Tabular Data Visualization 29 Jan 2024 · 0 repositories · arXiv:2402.07909
-
ReGAL: Refactoring Programs to Discover Generalizable Abstractions 29 Jan 2024 · 1 repository · arXiv:2401.16467Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Response Generation for Cognitive Behavioral Therapy with Large Language Models: Comparative Study with Socratic Questioning 29 Jan 2024 · 0 repositories · arXiv:2401.15966
-
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors 29 Jan 2024 · 0 repositories · arXiv:2401.16310
-
SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design 29 Jan 2024 · 1 repository · arXiv:2401.16456Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Spatio-Temporal Attention Graph Neural Network for Remaining Useful Life Prediction 29 Jan 2024 · 0 repositories · arXiv:2401.15964
-
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks 29 Jan 2024 · 1 repository · arXiv:2401.16589
-
TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting 29 Jan 2024 · 0 repositories · arXiv:2402.00066
-
Validation, Robustness, and Accuracy of Perturbation-Based Sensitivity Analysis Methods for Time-Series Deep Learning Models 29 Jan 2024 · 0 repositories · arXiv:2401.16521
-
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings 28 Jan 2024 · 1 repository · arXiv:2401.15713
-
Identifying and Improving Disability Bias in GPT-Based Resume Screening 28 Jan 2024 · 0 repositories · arXiv:2402.01732
-
PRE: A Peer Review Based Large Language Model Evaluator 28 Jan 2024 · 0 repositories · arXiv:2401.15641
-
SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection 28 Jan 2024 · 1 repository · arXiv:2401.15583Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts 28 Jan 2024 · 0 repositories · arXiv:2401.15798
-
A New Method for Vehicle Logo Recognition Based on Swin Transformer 27 Jan 2024 · 0 repositories · arXiv:2401.15458
-
Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization 27 Jan 2024 · 0 repositories · arXiv:2401.15496
-
ConvoSense: Overcoming Monotonous Commonsense Inferences for Conversational AI 27 Jan 2024 · 1 repository · arXiv:2401.15471
-
DataFrame QA: A Universal LLM Framework on DataFrame Question Answering Without Data Exposure 27 Jan 2024 · 0 repositories · arXiv:2401.15463