Methods › General › Feedforward Networks › Position-Wise Feed-Forward Layer › Papers, page 5
Position-Wise Feed-Forward Layer
Papers archive 2025-07-28
archive papers tagged: 13,895 · with a code link: 6,514 · where Syntology ran a sample: 2,229 (1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,229 of 13,895 tagged: 1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument)
Page 5 of 139: papers 401 to 500 of 13,895, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SatelliteCalculator: A Multi-Task Vision Foundation Model for Quantitative Remote Sensing Inversion 18 Apr 2025 · 0 repositories · arXiv:2504.13442
-
Towards Accurate and Interpretable Neuroblastoma Diagnosis via Contrastive Multi-scale Pathological Image Analysis 18 Apr 2025 · 1 repository · arXiv:2504.13754
-
Towards Scale-Aware Low-Light Enhancement via Structure-Guided Transformer Design 18 Apr 2025 · 1 repository · arXiv:2504.14075
-
Transformer Encoder and Multi-features Time2Vec for Financial Prediction 18 Apr 2025 · 0 repositories · arXiv:2504.13801
-
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective 18 Apr 2025 · 0 repositories · arXiv:2504.13558
-
Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models 17 Apr 2025 · 0 repositories · arXiv:2504.13068
-
Exploring Expert Failures Improves LLM Agent Tuning 17 Apr 2025 · 0 repositories · arXiv:2504.13145
-
Plain Transformers Can be Powerful Graph Learners 17 Apr 2025 · 0 repositories · arXiv:2504.12588
-
SSTAF: Spatial-Spectral-Temporal Attention Fusion Transformer for Motor Imagery Classification 17 Apr 2025 · 0 repositories · arXiv:2504.13220
-
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation 16 Apr 2025 · 0 repositories · arXiv:2504.11942
-
Approximation Bounds for Transformer Networks with Application to Regression 16 Apr 2025 · 0 repositories · arXiv:2504.12175
-
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts 16 Apr 2025 · 1 repository · arXiv:2504.12463Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Geometric Generality of Transformer-Based Gröbner Basis Computation 16 Apr 2025 · 1 repository · arXiv:2504.12465
-
GT-SVQ: A Linear-Time Graph Transformer for Node Classification Using Spiking Vector Quantization 16 Apr 2025 · 1 repository · arXiv:2504.11840
-
Human Aligned Compression for Robust Models 16 Apr 2025 · 1 repository · arXiv:2504.12255
-
Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study 16 Apr 2025 · 0 repositories · arXiv:2504.12422
-
Uncertainty-Guided Coarse-to-Fine Tumor Segmentation with Anatomy-Aware Post-Processing 16 Apr 2025 · 0 repositories · arXiv:2504.12215
-
Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach 16 Apr 2025 · 1 repository · arXiv:2504.11922
-
A Decade of Wheat Mapping for Lebanon 15 Apr 2025 · 0 repositories · arXiv:2504.11366
-
AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images 15 Apr 2025 · 1 repository · arXiv:2504.10972
-
Bridging Distribution Gaps in Time Series Foundation Model Pretraining with Prototype-Guided Normalization 15 Apr 2025 · 0 repositories · arXiv:2504.10900
-
Deep Learning-based Bathymetry Retrieval without In-situ Depths using Remote Sensing Imagery and SfM-MVS DSMs with Data Gaps 15 Apr 2025 · 1 repository · arXiv:2504.11416
-
Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification 15 Apr 2025 · 0 repositories · arXiv:2504.10916
-
Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs 15 Apr 2025 · 0 repositories · arXiv:2504.10982
-
Leveraging Point Transformers for Detecting Anatomical Landmarks in Digital Dentistry 15 Apr 2025 · 0 repositories · arXiv:2504.11418
-
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators 15 Apr 2025 · 0 repositories · arXiv:2504.10845
-
Multi-scale convolutional transformer network for motor imagery brain-computer interface 15 Apr 2025 · 1 repository
-
Progressive Rock Music Classification 15 Apr 2025 · 0 repositories · arXiv:2504.10821
-
Towards A Universal Graph Structural Encoder 15 Apr 2025 · 0 repositories · arXiv:2504.10917
-
Transformer-Based Model for Cold Start Mitigation in FaaS Architecture 15 Apr 2025 · 0 repositories · arXiv:2504.11338
-
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers 15 Apr 2025 · 0 repositories · arXiv:2504.11227
-
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models 14 Apr 2025 · 0 repositories · arXiv:2504.10615
-
Can LLMs handle WebShell detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework 14 Apr 2025 · 0 repositories · arXiv:2504.13811
-
EMAFusion: A Self-Optimizing System for Seamless LLM Selection and Integration 14 Apr 2025 · 0 repositories · arXiv:2504.10681
-
Global and Local Mamba Network for Multi-Modality Medical Image Super-Resolution 14 Apr 2025 · 0 repositories · arXiv:2504.10105
-
Integrating Vision and Location with Transformers: A Multimodal Deep Learning Framework for Medical Wound Analysis 14 Apr 2025 · 0 repositories · arXiv:2504.10452
-
Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content 14 Apr 2025 · 0 repositories · arXiv:2504.10679
-
Multi-Object Grounding via Hierarchical Contrastive Siamese Transformers 14 Apr 2025 · 0 repositories · arXiv:2504.10048
-
Multimodal Long Video Modeling Based on Temporal Dynamic Context 14 Apr 2025 · 1 repository · arXiv:2504.10443
-
Self-Controlled Dynamic Expansion Model for Continual Learning 14 Apr 2025 · 0 repositories · arXiv:2504.10561
-
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning 14 Apr 2025 · 1 repository · arXiv:2504.11493
-
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model 13 Apr 2025 · 1 repository · arXiv:2504.09421
-
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers 13 Apr 2025 · 0 repositories · arXiv:2504.09381
-
Enhanced Filterless Multi-Color VLC via QCT 13 Apr 2025 · 0 repositories · arXiv:2504.09743
-
Ensemble-Enhanced Graph Autoencoder with GAT and Transformer-Based Encoders for Robust Fault Diagnosis 13 Apr 2025 · 0 repositories · arXiv:2504.09427
-
Integrating Large Language Models for Automated Structural Analysis 13 Apr 2025 · 0 repositories · arXiv:2504.09754
-
Iterative Self-Training for Code Generation via Reinforced Re-Ranking 13 Apr 2025 · 0 repositories · arXiv:2504.09643
-
Trajectory-guided Motion Perception for Facial Expression Quality Assessment in Neurological Disorders 13 Apr 2025 · 1 repository · arXiv:2504.09530
-
Accurate Diagnosis of Respiratory Viruses Using an Explainable Machine Learning with Mid-Infrared Biomolecular Fingerprinting of Nasopharyngeal Secretions 12 Apr 2025 · 0 repositories · arXiv:2504.09211
-
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis 12 Apr 2025 · 0 repositories · arXiv:2504.09225
-
Multi-Modal Brain Tumor Segmentation via 3D Multi-Scale Self-attention and Cross-attention 12 Apr 2025 · 0 repositories · arXiv:2504.09088
-
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition 12 Apr 2025 · 0 repositories · arXiv:2504.09215
-
NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed Graph 12 Apr 2025 · 1 repository · arXiv:2504.09260
-
Adaptive Additive Parameter Updates of Vision Transformers for Few-Shot Continual Learning 11 Apr 2025 · 0 repositories · arXiv:2504.08982
-
DreamFuse: Adaptive Image Fusion with Diffusion Transformer 11 Apr 2025 · 0 repositories · arXiv:2504.08291
-
DrivAer Transformer: A high-precision and fast prediction method for vehicle aerodynamic drag coefficient based on the DrivAerNet++ dataset 11 Apr 2025 · 0 repositories · arXiv:2504.08217
-
Hypergraph Vision Transformers: Images are More than Nodes, More than Edges 11 Apr 2025 · 0 repositories · arXiv:2504.08710
-
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media 11 Apr 2025 · 0 repositories · arXiv:2504.12325
-
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner 11 Apr 2025 · 0 repositories · arXiv:2504.08247
-
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft 11 Apr 2025 · 0 repositories · arXiv:2504.08388
-
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization 11 Apr 2025 · 0 repositories · arXiv:2504.08398
-
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation 11 Apr 2025 · 0 repositories · arXiv:2504.08862
-
SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data 11 Apr 2025 · 0 repositories · arXiv:2504.08441
-
SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling 11 Apr 2025 · 0 repositories · arXiv:2504.08719
-
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering 11 Apr 2025 · 0 repositories · arXiv:2504.08269
-
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration 11 Apr 2025 · 0 repositories · arXiv:2504.08591
-
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks 10 Apr 2025 · 0 repositories · arXiv:2504.12321
-
Beating Transformers using Synthetic Cognition 10 Apr 2025 · 0 repositories · arXiv:2504.07619
-
Beyond Feature Importance: Feature Interactions in Predicting Post-Stroke Rigidity with Graph Explainable AI 10 Apr 2025 · 0 repositories · arXiv:2504.08150
-
Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition 10 Apr 2025 · 0 repositories · arXiv:2504.07792
-
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather 10 Apr 2025 · 1 repository · arXiv:2504.07625
-
Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation 10 Apr 2025 · 0 repositories · arXiv:2504.07691
-
Genetic Programming with Reinforcement Learning Trained Transformer for Real-World Dynamic Scheduling Problems 10 Apr 2025 · 0 repositories · arXiv:2504.07779
-
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability 10 Apr 2025 · 0 repositories · arXiv:2504.12320
-
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases 10 Apr 2025 · 1 repository · arXiv:2504.07606
-
JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture 10 Apr 2025 · 0 repositories · arXiv:2504.10512
-
Novel Pooling-based VGG-Lite for Pneumonia and Covid-19 Detection from Imbalanced Chest X-Ray Datasets 10 Apr 2025 · 0 repositories · arXiv:2504.07468
-
On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction 10 Apr 2025 · 0 repositories · arXiv:2504.08169
-
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs 10 Apr 2025 · 0 repositories · arXiv:2504.07866
-
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability 10 Apr 2025 · 0 repositories · arXiv:2504.07416Syntology 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction 10 Apr 2025 · 0 repositories · arXiv:2504.07357
-
Synthetic Fluency: Hallucinations, Confabulations, and the Creation of Irish Words in LLM-Generated Translations 10 Apr 2025 · 0 repositories · arXiv:2504.07680
-
AMAD: AutoMasked Attention for Unsupervised Multivariate Time Series Anomaly Detection 9 Apr 2025 · 0 repositories · arXiv:2504.06643
-
DyDiT++: Dynamic Diffusion Transformers for Efficient Visual Generation 9 Apr 2025 · 1 repository · arXiv:2504.06803
-
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation 9 Apr 2025 · 0 repositories · arXiv:2504.08806
-
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 9 Apr 2025 · 0 repositories · arXiv:2504.07083
-
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation 9 Apr 2025 · 1 repository · arXiv:2504.07072
-
Linguistic Interpretability of Transformer-based Language Models: a systematic review 9 Apr 2025 · 1 repository · arXiv:2504.08001
-
Assessing how hyperparameters impact Large Language Models' sarcasm detection performance 8 Apr 2025 · 0 repositories · arXiv:2504.06166
-
Fusing Global and Local: Transformer-CNN Synergy for Next-Gen Current Estimation 8 Apr 2025 · 0 repositories · arXiv:2504.07996
-
Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification 8 Apr 2025 · 1 repository · arXiv:2504.05583
-
Leveraging Auto-Distillation and Generative Self-Supervised Learning in Residual Graph Transformers for Enhanced Recommender Systems 8 Apr 2025 · 0 repositories · arXiv:2504.10500
-
Rethinking the Nested U-Net Approach: Enhancing Biomarker Segmentation with Attention Mechanisms and Multiscale Feature Fusion 8 Apr 2025 · 1 repository · arXiv:2504.06158
-
Boundary representation learning via Transformer 7 Apr 2025 · 0 repositories · arXiv:2504.07134
-
Content-Aware Transformer for All-in-one Image Restoration 7 Apr 2025 · 1 repository · arXiv:2504.04869
-
InstructionBench: An Instructional Video Understanding Benchmark 7 Apr 2025 · 0 repositories · arXiv:2504.05040
-
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision 7 Apr 2025 · 0 repositories · arXiv:2504.04903
-
OmniEcon Nexus: Global Microeconomic Simulation Engine 7 Apr 2025 · 1 repository
-
One-Minute Video Generation with Test-Time Training 7 Apr 2025 · 0 repositories · arXiv:2504.05298
-
One Quantizer is Enough: Toward a Lightweight Audio Codec 7 Apr 2025 · 1 repository · arXiv:2504.04949