Methods › General › Attention Mechanisms › Attention › Papers, page 201
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 201 of 316: papers 20,001 to 20,100 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Efficient Large-Scale Visual Representation Learning And Evaluation 22 May 2023 · 0 repositories · arXiv:2305.13399
-
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study 22 May 2023 · 1 repository · arXiv:2305.13062
-
ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist Examination 22 May 2023 · 1 repository · arXiv:2305.12945
-
Exploring Energy-based Language Models with Different Architectures and Training Methods for Speech Recognition 22 May 2023 · 2 repositories · arXiv:2305.12676
-
Flover: A Temporal Fusion Framework for Efficient Autoregressive Model Parallel Inference 22 May 2023 · 1 repository · arXiv:2305.13484
-
G3Detector: General GPT-Generated Text Detector 22 May 2023 · 0 repositories · arXiv:2305.12680
-
GATology for Linguistics: What Syntactic Dependencies It Knows 22 May 2023 · 0 repositories · arXiv:2305.13403
-
GNCformer Enhanced Self-attention for Automatic Speech Recognition 22 May 2023 · 0 repositories · arXiv:2305.12755
-
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints 22 May 2023 · 4 repositories · arXiv:2305.13245Syntology 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
How Language Model Hallucinations Can Snowball 22 May 2023 · 1 repository · arXiv:2305.13534
-
SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two Perspectives 22 May 2023 · 0 repositories · arXiv:2305.13192
-
InheritSumm: A General, Versatile and Compact Summarizer by Distilling from GPT 22 May 2023 · 0 repositories · arXiv:2305.13083
-
Iterative Forward Tuning Boosts In-Context Learning in Language Models 22 May 2023 · 1 repository · arXiv:2305.13016
-
Language-Agnostic Bias Detection in Language Models with Bias Probing 22 May 2023 · 1 repository · arXiv:2305.13302
-
Learning Pedestrian Actions to Ensure Safe Autonomous Driving 22 May 2023 · 0 repositories · arXiv:2305.13051
-
Learning Subpocket Prototypes for Generalizable Structure-based Drug Design 22 May 2023 · 1 repository · arXiv:2305.13997Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities 22 May 2023 · 1 repository · arXiv:2305.13168Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Atomic Inference for NLI with Generated Facts as Atoms 22 May 2023 · 1 repository · arXiv:2305.13214
-
MAILEX: Email Event and Argument Extraction 22 May 2023 · 1 repository · arXiv:2305.13469
-
Materialistic: Selecting Similar Materials in Images 22 May 2023 · 0 repositories · arXiv:2305.13291
-
Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations 22 May 2023 · 1 repository · arXiv:2305.13299
-
Multi-Task Instruction Tuning of LLaMa for Specific Scenarios: A Preliminary Study on Writing Assistance 22 May 2023 · 0 repositories · arXiv:2305.13225
-
Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design 22 May 2023 · 0 repositories · arXiv:2305.13297
-
RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text 22 May 2023 · 2 repositories · arXiv:2305.13304Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
RWKV: Reinventing RNNs for the Transformer Era 22 May 2023 · 14 repositories · arXiv:2305.13048Syntology official (archive's flag): 3 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables 22 May 2023 · 1 repository · arXiv:2305.13186
-
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations 22 May 2023 · 1 repository · arXiv:2305.13235
-
Spatiotemporal Attention-based Semantic Compression for Real-time Video Recognition 22 May 2023 · 0 repositories · arXiv:2305.12796
-
Stock and market index prediction using Informer network 22 May 2023 · 0 repositories · arXiv:2305.14382
-
Syntactic Knowledge via Graph Attention with BERT in Machine Translation 22 May 2023 · 0 repositories · arXiv:2305.13413
-
The Emergence of Economic Rationality of GPT 22 May 2023 · 0 repositories · arXiv:2305.12763
-
Tokenized Graph Transformer with Neighborhood Augmentation for Node Classification in Large Graphs 22 May 2023 · 0 repositories · arXiv:2305.12677
-
Investigating Agency of LLMs in Human-AI Collaboration Tasks 22 May 2023 · 0 repositories · arXiv:2305.12815
-
TSPTQ-ViT: Two-scaled post-training quantization for vision transformer 22 May 2023 · 0 repositories · arXiv:2305.12901
-
VDT: General-purpose Video Diffusion Transformers via Mask Modeling 22 May 2023 · 1 repository · arXiv:2305.13311Syntology official (archive's flag): 4 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 2 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
VideoLLM: Modeling Video Sequence with Large Language Models 22 May 2023 · 1 repository · arXiv:2305.13292
-
Why current rain denoising models fail on CycleGAN created rain images in autonomous driving 22 May 2023 · 0 repositories · arXiv:2305.12983
-
A Deeper (Autoregressive) Approach to Non-Convergent Discourse Parsing 21 May 2023 · 0 repositories · arXiv:2305.12510
-
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers 21 May 2023 · 0 repositories · arXiv:2305.12563
-
BertRLFuzzer: A BERT and Reinforcement Learning Based Fuzzer 21 May 2023 · 1 repository · arXiv:2305.12534
-
Bi-ViT: Pushing the Limit of Vision Transformer Quantization 21 May 2023 · 0 repositories · arXiv:2305.12354
-
BiasAsker: Measuring the Bias in Conversational AI System 21 May 2023 · 1 repository · arXiv:2305.12434
-
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning 21 May 2023 · 1 repository · arXiv:2305.12599
-
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 21 May 2023 · 1 repository · arXiv:2305.12474Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Explaining How Transformers Use Context to Build Predictions 21 May 2023 · 1 repository · arXiv:2305.12535Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
F-PABEE: Flexible-patience-based Early Exiting for Single-label and Multi-label text Classification Tasks 21 May 2023 · 0 repositories · arXiv:2305.11916
-
Gene Set Summarization using Large Language Models 21 May 2023 · 1 repository · arXiv:2305.13338Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GPT-3.5, GPT-4, or BARD? Evaluating LLMs Reasoning Ability in Zero-Shot Setting and Performance Boosting Through Prompts 21 May 2023 · 0 repositories · arXiv:2305.12477
-
DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection 21 May 2023 · 0 repositories · arXiv:2305.12519
-
HIINT: Historical, Intra- and Inter- personal Dynamics Modeling with Cross-person Memory Transformer 21 May 2023 · 0 repositories · arXiv:2305.12369
-
Infor-Coef: Information Bottleneck-based Dynamic Token Downsampling for Compact and Efficient language model 21 May 2023 · 0 repositories · arXiv:2305.12458
-
IR Models and the COVID-19 Pandemic: A Comparative Study of Performance and Challenges 21 May 2023 · 0 repositories · arXiv:2305.12528
-
Learning Joint 2D & 3D Diffusion Models for Complete Molecule Generation 21 May 2023 · 2 repositories · arXiv:2305.12347Syntology official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers 21 May 2023 · 1 repository · arXiv:2305.12567
-
WiFi-TCN: Temporal Convolution for Human Interaction Recognition based on WiFi signal 21 May 2023 · 0 repositories · arXiv:2305.18211
-
Teaching the Pre-trained Model to Generate Simple Texts for Text Simplification 21 May 2023 · 1 repository · arXiv:2305.12463
-
Temporal Fusion Transformers for Streamflow Prediction: Value of Combining Attention with Recurrence 21 May 2023 · 0 repositories · arXiv:2305.12335
-
TheoremQA: A Theorem-driven Question Answering dataset 21 May 2023 · 1 repository · arXiv:2305.12524Syntology official (archive's flag): 1 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Autoregressive Modeling with Lookahead Attention 20 May 2023 · 0 repositories · arXiv:2305.12272
-
Can NLP Models Correctly Reason Over Contexts that Break the Common Assumptions? 20 May 2023 · 0 repositories · arXiv:2305.12096
-
CDJUR-BR -- A Golden Collection of Legal Document from Brazilian Justice with Fine-Grained Named Entities 20 May 2023 · 0 repositories · arXiv:2305.18315
-
Comparative Analysis of Deep Learning Models for Brand Logo Classification in Real-World Scenarios 20 May 2023 · 0 repositories · arXiv:2305.12242
-
Contextualizing Argument Quality Assessment with Relevant Knowledge 20 May 2023 · 1 repository · arXiv:2305.12280
-
Experimental results from applying GPT-4 to an unpublished formal language 20 May 2023 · 0 repositories · arXiv:2305.12196
-
Learning to Compose Representations of Different Encoder Layers towards Improving Compositional Generalization 20 May 2023 · 0 repositories · arXiv:2305.12169
-
LogiCoT: Logical Chain-of-Thought Instruction-Tuning 20 May 2023 · 1 repository · arXiv:2305.12147
-
CARD: Channel Aligned Robust Blend Transformer for Time Series Forecasting 20 May 2023 · 1 repository · arXiv:2305.12095Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Practical PCG Through Large Language Models 20 May 2023 · 0 repositories · arXiv:2305.18243
-
Revisiting the Architectures like Pointer Networks to Efficiently Improve the Next Word Distribution, Summarization Factuality, and Beyond 20 May 2023 · 1 repository · arXiv:2305.12289
-
SEntFiN 1.0: Entity-Aware Sentiment Analysis for Financial News 20 May 2023 · 0 repositories · arXiv:2305.12257
-
What Makes for Good Visual Tokenizers for Large Language Models? 20 May 2023 · 1 repository · arXiv:2305.12223Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
A Sequence-to-Sequence Approach for Arabic Pronoun Resolution 19 May 2023 · 0 repositories · arXiv:2305.11529
-
AutoTrial: Prompting Language Models for Clinical Trial Design 19 May 2023 · 0 repositories · arXiv:2305.11366
-
Brain Captioning: Decoding human brain activity into images and text 19 May 2023 · 1 repository · arXiv:2305.11560
-
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search 19 May 2023 · 0 repositories · arXiv:2305.11626
-
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding 19 May 2023 · 2 repositories · arXiv:2305.12031Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate 19 May 2023 · 1 repository · arXiv:2305.11595Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning 19 May 2023 · 0 repositories · arXiv:2305.11383
-
Efficient Mixed Transformer for Single Image Super-Resolution 19 May 2023 · 2 repositories · arXiv:2305.11403
-
Enhancing Short-Term Wind Speed Forecasting using Graph Attention and Frequency-Enhanced Mechanisms 19 May 2023 · 0 repositories · arXiv:2305.11526
-
Exploring the Upper Limits of Text-Based Collaborative Filtering Using Large Language Models: Discoveries and Insights 19 May 2023 · 0 repositories · arXiv:2305.11700
-
Eye-SpatialNet: Spatial Information Extraction from Ophthalmology Notes 19 May 2023 · 0 repositories · arXiv:2305.11948
-
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models 19 May 2023 · 0 repositories · arXiv:2305.11414
-
Graph Propagation Transformer for Graph Representation Learning 19 May 2023 · 1 repository · arXiv:2305.11424Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 8 harvested samples) · 8 pointer-only (licence)
-
How Does Generative Retrieval Scale to Millions of Passages? 19 May 2023 · 0 repositories · arXiv:2305.11841
-
Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition 19 May 2023 · 0 repositories · arXiv:2305.11569
-
Multimodal Web Navigation with Instruction-Finetuned Foundation Models 19 May 2023 · 0 repositories · arXiv:2305.11854
-
OL-Transformer: A Fast and Universal Surrogate Simulator for Optical Multilayer Thin Film Structures 19 May 2023 · 0 repositories · arXiv:2305.11984
-
PointGPT: Auto-regressively Generative Pre-training from Point Clouds 19 May 2023 · 1 repository · arXiv:2305.11487Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Reciprocal Attention Mixing Transformer for Lightweight Image Restoration 19 May 2023 · 1 repository · arXiv:2305.11474
-
Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation 19 May 2023 · 1 repository · arXiv:2305.11685
-
Reducing Sequence Length by Predicting Edit Operations with Large Language Models 19 May 2023 · 0 repositories · arXiv:2305.11862
-
Scaling laws for language encoding models in fMRI 19 May 2023 · 1 repository · arXiv:2305.11863
-
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models 19 May 2023 · 1 repository · arXiv:2305.11840
-
Self-Agreement: A Framework for Fine-tuning Language Models to Find Agreement among Diverse Opinions 19 May 2023 · 0 repositories · arXiv:2305.11460
-
Self-QA: Unsupervised Knowledge Guided Language Model Alignment 19 May 2023 · 1 repository · arXiv:2305.11952
-
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs 19 May 2023 · 0 repositories · arXiv:2305.11461
-
Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery 19 May 2023 · 2 repositories · arXiv:2305.11692Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks 18 May 2023 · 2 repositories · arXiv:2305.11073
-
Ahead-of-Time P-Tuning 18 May 2023 · 0 repositories · arXiv:2305.10835