Methods › General › Feedforward Networks › Position-Wise Feed-Forward Layer › Papers, page 12
Position-Wise Feed-Forward Layer
Papers archive 2025-07-28
archive papers tagged: 13,895 · with a code link: 6,514 · where Syntology ran a sample: 2,229 (1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,229 of 13,895 tagged: 1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument)
Page 12 of 139: papers 1,101 to 1,200 of 13,895, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
GNN-DT: Graph Neural Network Enhanced Decision Transformer for Efficient Optimization in Dynamic Environments 3 Feb 2025 · 1 repository · arXiv:2502.01778
-
Joint Localization and Activation Editing for Low-Resource Fine-Tuning 3 Feb 2025 · 1 repository · arXiv:2502.01179
-
Scalable Language Models with Posterior Inference of Latent Thought Vectors 3 Feb 2025 · 0 repositories · arXiv:2502.01567
-
Toward Neurosymbolic Program Comprehension 3 Feb 2025 · 0 repositories · arXiv:2502.01806
-
Transformers trained on proteins can learn to attend to Euclidean distance 3 Feb 2025 · 1 repository · arXiv:2502.01533
-
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale 2 Feb 2025 · 1 repository · arXiv:2502.01681Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Estimating forest carbon stocks from high-resolution remote sensing imagery by reducing domain shift with style transfer 2 Feb 2025 · 0 repositories · arXiv:2502.00784
-
A Study on the Performance of U-Net Modifications in Retroperitoneal Tumor Segmentation 1 Feb 2025 · 1 repository · arXiv:2502.00314
-
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset 1 Feb 2025 · 0 repositories · arXiv:2502.01676
-
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer 1 Feb 2025 · 0 repositories · arXiv:2502.00571
-
Converting Transformers into DGNNs Form 1 Feb 2025 · 1 repository · arXiv:2502.00585
-
MambaGlue: Fast and Robust Local Feature Matching With Mamba 1 Feb 2025 · 1 repository · arXiv:2502.00462
-
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition 1 Feb 2025 · 1 repository · arXiv:2502.00547
-
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective 1 Feb 2025 · 0 repositories · arXiv:2502.00281
-
VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility 1 Feb 2025 · 1 repository · arXiv:2502.00543
-
Accelerating Diffusion Transformer via Error-Optimized Cache 31 Jan 2025 · 0 repositories · arXiv:2501.19243
-
CerraData-4MM: A multimodal benchmark dataset on Cerrado for land use and land cover classification 31 Jan 2025 · 1 repository · arXiv:2502.00083
-
ContextFormer: Redefining Efficiency in Semantic Segmentation 31 Jan 2025 · 0 repositories · arXiv:2501.19255
-
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game 31 Jan 2025 · 1 repository · arXiv:2501.19398
-
Homogeneity Bias as Differential Sampling Uncertainty in Language Models 31 Jan 2025 · 0 repositories · arXiv:2501.19337
-
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media 31 Jan 2025 · 0 repositories · arXiv:2502.01658
-
Privacy Preserving Charge Location Prediction for Electric Vehicles 31 Jan 2025 · 0 repositories · arXiv:2502.00068
-
Strassen Attention: Unlocking Compositional Abilities in Transformers Based on a New Lower Bound Method 31 Jan 2025 · 0 repositories · arXiv:2501.19215
-
A Unified Perspective on the Dynamics of Deep Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18322
-
Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18237
-
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights 30 Jan 2025 · 0 repositories · arXiv:2501.18596
-
Evaluating Large Language Models in Vulnerability Detection Under Variable Context Windows 30 Jan 2025 · 0 repositories · arXiv:2502.00064
-
GDformer: Going Beyond Subsequence Isolation for Multivariate Time Series Anomaly Detection 30 Jan 2025 · 1 repository · arXiv:2501.18196
-
GENIE: Generative Note Information Extraction model for structuring EHR data 30 Jan 2025 · 0 repositories · arXiv:2501.18435
-
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning 30 Jan 2025 · 0 repositories · arXiv:2501.18187
-
MatIR: A Hybrid Mamba-Transformer Image Restoration Model 30 Jan 2025 · 1 repository · arXiv:2501.18401
-
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer 30 Jan 2025 · 1 repository · arXiv:2501.18427
-
State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence 30 Jan 2025 · 0 repositories · arXiv:2501.18356
-
Survey and Improvement Strategies for Gene Prioritization with Large Language Models 30 Jan 2025 · 0 repositories · arXiv:2501.18794
-
Transformer Semantic Genetic Programming for Symbolic Regression 30 Jan 2025 · 0 repositories · arXiv:2501.18479
-
Unraveling the Capabilities of Language Models in News Summarization 30 Jan 2025 · 1 repository · arXiv:2501.18128
-
2SSP: A Two-Stage Framework for Structured Pruning of LLMs 29 Jan 2025 · 1 repository · arXiv:2501.17771
-
ContourFormer:Real-Time Contour-Based End-to-End Instance Segmentation Transformer 29 Jan 2025 · 1 repository · arXiv:2501.17688
-
DINT Transformer 29 Jan 2025 · 0 repositories · arXiv:2501.17486
-
Hybrid Graphs for Table-and-Text based Question Answering using LLMs 29 Jan 2025 · 0 repositories · arXiv:2501.17767
-
Leveraging In-Context Learning and Retrieval-Augmented Generation for Automatic Question Generation in Educational Domains 29 Jan 2025 · 0 repositories · arXiv:2501.17397
-
Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Model: An Initial Multi-layered Tabular Review 29 Jan 2025 · 0 repositories · arXiv:2501.18644
-
Shared DIFF Transformer 29 Jan 2025 · 0 repositories · arXiv:2501.17900
-
Transformer Based Time-Series Forecasting for Stock 29 Jan 2025 · 0 repositories · arXiv:2502.09625
-
TransRAD: Retentive Vision Transformer for Enhanced Radar Object Detection 29 Jan 2025 · 1 repository · arXiv:2501.17977
-
Watch Your STEPP: Semantic Traversability Estimation using Pose Projected Features 29 Jan 2025 · 0 repositories · arXiv:2501.17594
-
An Attention-Locating Algorithm for Eliminating Background Effects in Fine-grained Visual Classification 28 Jan 2025 · 1 repository
-
Chinese Stock Prediction Based on a Multi-Modal Transformer Framework: Macro-Micro Information Fusion 28 Jan 2025 · 0 repositories · arXiv:2501.16621
-
FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation 28 Jan 2025 · 0 repositories · arXiv:2501.16778
-
Generative quantum combinatorial optimization by means of a novel conditional generative quantum eigensolver 28 Jan 2025 · 0 repositories · arXiv:2501.16986
-
Graph Transformers for inverse physics: reconstructing flows around arbitrary 2D airfoils 28 Jan 2025 · 0 repositories · arXiv:2501.17081
-
JRE-L: Journalist, Reader, and Editor LLMs in the Loop for Science Journalism for the General Audience 28 Jan 2025 · 1 repository · arXiv:2501.16865
-
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models 28 Jan 2025 · 1 repository · arXiv:2501.17088Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition 28 Jan 2025 · 0 repositories · arXiv:2501.17011
-
Scenario Understanding of Traffic Scenes Through Large Visual Language Models 28 Jan 2025 · 0 repositories · arXiv:2501.17131
-
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification 28 Jan 2025 · 1 repository · arXiv:2501.17260
-
Cross-Domain Semantic Segmentation with Large Language Model-Assisted Descriptor Generation 27 Jan 2025 · 0 repositories · arXiv:2501.16467
-
LCTG Bench: LLM Controlled Text Generation Benchmark 27 Jan 2025 · 1 repository · arXiv:2501.15875
-
Object Detection for Medical Image Analysis: Insights from the RT-DETR Model 27 Jan 2025 · 0 repositories · arXiv:2501.16469
-
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer 27 Jan 2025 · 0 repositories · arXiv:2501.16227
-
Phase Transitions in Large Language Models and the O(N) Model 27 Jan 2025 · 0 repositories · arXiv:2501.16241
-
UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification 27 Jan 2025 · 0 repositories · arXiv:2501.16542
-
Adapting Biomedical Abstracts into Plain language using Large Language Models 26 Jan 2025 · 0 repositories · arXiv:2501.15700
-
AI-Driven Secure Data Sharing: A Trustworthy and Privacy-Preserving Approach 26 Jan 2025 · 0 repositories · arXiv:2501.15363
-
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer 26 Jan 2025 · 1 repository · arXiv:2501.15570
-
Classifying Deepfakes Using Swin Transformers 26 Jan 2025 · 0 repositories · arXiv:2501.15656
-
Decentralized Low-Rank Fine-Tuning of Large Language Models 26 Jan 2025 · 0 repositories · arXiv:2501.15361
-
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets 26 Jan 2025 · 0 repositories · arXiv:2501.15624
-
SedarEval: Automated Evaluation using Self-Adaptive Rubrics 26 Jan 2025 · 1 repository · arXiv:2501.15595Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel 26 Jan 2025 · 0 repositories · arXiv:2501.15665
-
Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment 26 Jan 2025 · 0 repositories · arXiv:2501.16394
-
An AI-Driven Live Systematic Reviews in the Brain-Heart Interconnectome: Minimizing Research Waste and Advancing Evidence Synthesis 25 Jan 2025 · 1 repository · arXiv:2501.17181
-
ILETIA: An AI-enhanced method for individualized trigger-oocyte pickup interval estimation of progestin-primed ovarian stimulation protocol 25 Jan 2025 · 0 repositories · arXiv:2501.16386
-
Knowledge Hierarchy Guided Biological-Medical Dataset Distillation for Domain LLM Training 25 Jan 2025 · 0 repositories · arXiv:2501.15108
-
LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering 25 Jan 2025 · 0 repositories · arXiv:2501.17183
-
TranStable: Towards Robust Pixel-level Online Video Stabilization by Jointing Transformer and CNN 25 Jan 2025 · 0 repositories · arXiv:2501.15138
-
Using Large Language Models for education managements in Vietnamese with low resources 25 Jan 2025 · 0 repositories · arXiv:2501.15022
-
Advances in Set Function Learning: A Survey of Techniques and Applications 24 Jan 2025 · 0 repositories · arXiv:2501.14991
-
Automatic detection and prediction of nAMD activity change in retinal OCT using Siamese networks and Wasserstein Distance for ordinality 24 Jan 2025 · 1 repository · arXiv:2501.14323
-
Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models 24 Jan 2025 · 0 repositories · arXiv:2501.14273
-
Low-rank Prompt Interaction for Continual Vision-Language Retrieval 24 Jan 2025 · 1 repository · arXiv:2501.14369
-
On the locality bias and results in the Long Range Arena 24 Jan 2025 · 0 repositories · arXiv:2501.14850
-
Rethinking Table Instruction Tuning 24 Jan 2025 · 1 repository · arXiv:2501.14693
-
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation 24 Jan 2025 · 0 repositories · arXiv:2501.14679
-
Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet Extraction 24 Jan 2025 · 0 repositories · arXiv:2501.14144
-
UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis 24 Jan 2025 · 0 repositories · arXiv:2501.16380
-
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention 24 Jan 2025 · 0 repositories · arXiv:2501.14577
-
5G LDPC Linear Transformer for Channel Decoding 23 Jan 2025 · 1 repository · arXiv:2501.14102
-
An Efficient Diffusion-based Non-Autoregressive Solver for Traveling Salesman Problem 23 Jan 2025 · 1 repository · arXiv:2501.13767Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Document-Level Sentiment Analysis of Urdu Text Using Deep Learning Techniques 23 Jan 2025 · 0 repositories · arXiv:2501.17175
-
EgoHand: Ego-centric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMUs 23 Jan 2025 · 1 repository · arXiv:2501.13805
-
Enhancing Biomedical Relation Extraction with Directionality 23 Jan 2025 · 1 repository · arXiv:2501.14079
-
Ensuring Medical AI Safety: Explainable AI-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data 23 Jan 2025 · 1 repository · arXiv:2501.13818
-
FreEformer: Frequency Enhanced Transformer for Multivariate Time Series Forecasting 23 Jan 2025 · 1 repository · arXiv:2501.13989Syntology official (archive's flag): 7 ran · 7 ran (of which 7 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 7 samples that ran constructed an object rather than computing a result (of 10 harvested samples) · 10 pointer-only (licence)
-
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language 23 Jan 2025 · 0 repositories · arXiv:2501.14073
-
LLMs Can Plan Only If We Tell Them 23 Jan 2025 · 0 repositories · arXiv:2501.13545
-
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods 23 Jan 2025 · 0 repositories · arXiv:2501.13484
-
ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection 23 Jan 2025 · 1 repository · arXiv:2501.14004
-
Multi-Level Attention and Contrastive Learning for Enhanced Text Classification with an Optimized Transformer 23 Jan 2025 · 0 repositories · arXiv:2501.13467
-
Polyhedra Encoding Transformers: Enhancing Diffusion MRI Analysis Beyond Voxel and Volumetric Embedding 23 Jan 2025 · 0 repositories · arXiv:2501.13352