{"url":"/method/graph-transformer","slug":"graph-transformer","name":"Graph Transformer","full_name":"Graph Transformer","full_name_withheld":false,"description_markdown":"This is **Graph Transformer** method, proposed as a generalization of [Transformer](https://paperswithcode.com/method/transformer) Neural Network architectures, for arbitrary graphs.\r\n\r\nCompared to the original Transformer, the highlights of the presented architecture are:\r\n\r\n- The attention mechanism is a function of neighborhood connectivity for each node in the graph.  \r\n- The position encoding is represented by Laplacian eigenvectors, which naturally generalize the sinusoidal positional encodings often used in NLP.  \r\n- The [layer normalization](https://paperswithcode.com/method/layer-normalization) is replaced by a [batch normalization](https://paperswithcode.com/method/batch-normalization) layer.  \r\n- The architecture is extended to have edge representation, which can be critical to tasks with rich information on the edges, or pairwise interactions (such as bond types in molecules, or relationship type in KGs. etc).","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2012.09699v2","title":"A Generalization of Transformer Networks to Graphs","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/graphdeeplearning/graphtransformer","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Graphs","area_id":"graphs","collection":"Graph Models","url":"/methods/category/graph-models","pwc_aliases":[]}],"n_papers_tagged":298,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"H$^2$GFM: Towards unifying Homogeneity and Heterogeneity on Text-Attributed Graphs","date":"2025-06-10","arxiv_id":"2506.08298","n_code_links":0,"syntology":null},{"paper":null,"title":"HGFormer: A Hierarchical Graph Transformer Framework for Two-Stage Colonel Blotto Games via Reinforcement Learning","date":"2025-06-10","arxiv_id":"2506.08580","n_code_links":0,"syntology":null},{"paper":null,"title":"Quantum Graph Transformer for NLP Sentiment Classification","date":"2025-06-09","arxiv_id":"2506.07937","n_code_links":0,"syntology":null},{"paper":"/paper/opengt-a-comprehensive-benchmark-for-graph","title":"OpenGT: A Comprehensive Benchmark For Graph Transformers","date":"2025-06-05","arxiv_id":"2506.04765","n_code_links":1,"syntology":null},{"paper":"/paper/learning-pyramid-structured-long-range-1","title":"Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation","date":"2025-06-03","arxiv_id":"2506.02853","n_code_links":1,"syntology":null},{"paper":"/paper/2505-10960","title":"Relational Graph Transformer","date":"2025-05-16","arxiv_id":"2505.10960","n_code_links":1,"syntology":{"ran":1,"of":4,"unverified":3,"pointer_only":0}},{"paper":null,"title":"All You Need Is Synthetic Task Augmentation","date":"2025-05-15","arxiv_id":"2505.10120","n_code_links":0,"syntology":null},{"paper":null,"title":"SAR-GTR: Attributed Scattering Information Guided SAR Graph Transformer Recognition Algorithm","date":"2025-05-13","arxiv_id":"2505.08547","n_code_links":0,"syntology":null},{"paper":"/paper/structural-temporal-coupling-anomaly","title":"Structural-Temporal Coupling Anomaly Detection with Dynamic Graph Transformer","date":"2025-05-13","arxiv_id":"2505.08330","n_code_links":1,"syntology":null},{"paper":"/paper/fused3s-fast-sparse-attention-on-tensor-cores","title":"Fused3S: Fast Sparse Attention on Tensor Cores","date":"2025-05-12","arxiv_id":"2505.08098","n_code_links":1,"syntology":null},{"paper":"/paper/mitigating-degree-bias-in-graph","title":"Mitigating Degree Bias in Graph Representation Learning with Learnable Structural Augmentation and Structural Self-Attention","date":"2025-04-21","arxiv_id":"2504.15075","n_code_links":1,"syntology":null},{"paper":null,"title":"HAECcity: Open-Vocabulary Scene Understanding of City-Scale Point Clouds with Superpoint Graph Clustering","date":"2025-04-18","arxiv_id":"2504.13590","n_code_links":0,"syntology":null},{"paper":"/paper/gt-svq-a-linear-time-graph-transformer-for","title":"GT-SVQ: A Linear-Time Graph Transformer for Node Classification Using Spiking Vector Quantization","date":"2025-04-16","arxiv_id":"2504.11840","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards A Universal Graph Structural Encoder","date":"2025-04-15","arxiv_id":"2504.10917","n_code_links":0,"syntology":null},{"paper":null,"title":"Ensemble-Enhanced Graph Autoencoder with GAT and Transformer-Based Encoders for Robust Fault Diagnosis","date":"2025-04-13","arxiv_id":"2504.09427","n_code_links":0,"syntology":null},{"paper":"/paper/nettag-a-multimodal-rtl-and-layout-aligned","title":"NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed Graph","date":"2025-04-12","arxiv_id":"2504.09260","n_code_links":1,"syntology":null},{"paper":null,"title":"Leveraging Auto-Distillation and Generative Self-Supervised Learning in Residual Graph Transformers for Enhanced Recommender Systems","date":"2025-04-08","arxiv_id":"2504.10500","n_code_links":0,"syntology":null},{"paper":null,"title":"Graphs are everywhere -- Psst! In Music Recommendation too","date":"2025-04-03","arxiv_id":"2504.02598","n_code_links":0,"syntology":null},{"paper":null,"title":"Graph Transformer-Based Flood Susceptibility Mapping: Application to the French Riviera and Railway Infrastructure Under Climate Change","date":"2025-03-31","arxiv_id":"2504.03727","n_code_links":0,"syntology":null},{"paper":null,"title":"TacticExpert: Spatial-Temporal Graph Language Model for Basketball Tactics","date":"2025-03-13","arxiv_id":"2503.10722","n_code_links":0,"syntology":null},{"paper":null,"title":"FMCHS: Advancing Traditional Chinese Medicine Herb Recommendation with Fusion of Multiscale Correlations of Herbs and Symptoms","date":"2025-03-07","arxiv_id":"2503.05167","n_code_links":0,"syntology":null},{"paper":"/paper/graph-transformer-with-disease-subgraph","title":"Graph Transformer with Disease Subgraph Positional Encoding for Improved Comorbidity Prediction","date":"2025-03-04","arxiv_id":"2503.03046","n_code_links":1,"syntology":null},{"paper":"/paper/co-mtp-a-cooperative-trajectory-prediction","title":"Co-MTP: A Cooperative Trajectory Prediction Framework with Multi-Temporal Fusion for Autonomous Driving","date":"2025-02-23","arxiv_id":"2502.16589","n_code_links":1,"syntology":null},{"paper":"/paper/lightweight-yet-efficient-an-external","title":"Lightweight yet Efficient: An External Attentive Graph Convolutional Network with Positional Prompts for Sequential Recommendation","date":"2025-02-21","arxiv_id":"2502.15331","n_code_links":1,"syntology":null},{"paper":null,"title":"Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning","date":"2025-02-19","arxiv_id":"2502.13754","n_code_links":0,"syntology":null},{"paper":null,"title":"GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion","date":"2025-02-17","arxiv_id":"2502.11471","n_code_links":0,"syntology":null},{"paper":"/paper/towards-mechanistic-interpretability-of-graph","title":"Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs","date":"2025-02-17","arxiv_id":"2502.12352","n_code_links":1,"syntology":null},{"paper":"/paper/biologically-plausible-brain-graph","title":"Biologically Plausible Brain Graph Transformer","date":"2025-02-13","arxiv_id":"2502.08958","n_code_links":1,"syntology":{"ran":3,"of":8,"unverified":5,"pointer_only":8}},{"paper":null,"title":"Deep Semantic Graph Learning via LLM based Node Enhancement","date":"2025-02-11","arxiv_id":"2502.07982","n_code_links":0,"syntology":null},{"paper":"/paper/medgnn-towards-multi-resolution","title":"MedGNN: Towards Multi-resolution Spatiotemporal Graph Learning for Medical Time Series Classification","date":"2025-02-06","arxiv_id":"2502.04515","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/representation-learning","name":"Representation Learning","papers":54},{"task":"/task/node-classification","name":"Node Classification","papers":38},{"task":"/task/graph-learning","name":"Graph Learning","papers":37},{"task":"/task/graph-representation-learning","name":"Graph Representation Learning","papers":28},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":24},{"task":"/task/prediction","name":"Prediction","papers":19},{"task":"/task/graph-regression","name":"Graph Regression","papers":18},{"task":"/task/graph-classification","name":"Graph Classification","papers":17},{"task":"/task/link-prediction","name":"Link Prediction","papers":17},{"task":"/task/property-prediction","name":"Property Prediction","papers":13},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":12},{"task":"/task/drug-discovery","name":"Drug Discovery","papers":12},{"task":"/task/graph-attention","name":"Graph Attention","papers":12},{"task":"/task/graph-generation","name":"Graph Generation","papers":12},{"task":"/task/molecular-property-prediction","name":"Molecular Property Prediction","papers":12},{"task":"/task/recommendation-systems","name":"Recommendation Systems","papers":11},{"task":"/task/denoising","name":"Denoising","papers":10},{"task":"/task/decoder","name":"Decoder","papers":9},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":9},{"task":"/task/language-modelling","name":"Language Modelling","papers":8}],"tasks_shown":20,"n_tasks":276,"usage_by_year":[{"year":"2020","papers":2},{"year":"2021","papers":26},{"year":"2022","papers":44},{"year":"2023","papers":87},{"year":"2024","papers":100},{"year":"2025","papers":39}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/graph-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}