{"url":"/method/set-transformer","slug":"set-transformer","name":"Set Transformer","full_name":"Set Transformer","full_name_withheld":false,"description_markdown":"Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set, models used to address them should be permutation invariant. We present an attention-based neural network module, the Set Transformer, specifically designed to model interactions among elements in the input set. The model consists of an encoder and a decoder, both of which rely on attention mechanisms. In an effort to reduce computational complexity, we introduce an attention scheme inspired by inducing point methods from sparse Gaussian process literature. It reduces the computation time of self-attention from quadratic to linear in the number of elements in the set. We show that our model is theoretically attractive and we evaluate it on a range of tasks, demonstrating the state-of-the-art performance compared to recent methods for set-structured data.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks","paper":"/paper/set-transformer-a-framework-for-attention","first_author":"Juho Lee","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/set-transformer-a-framework-for-attention"},"source":{"url":"https://arxiv.org/abs/1810.00825v3","title":"Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Attention Mechanisms","url":"/methods/category/attention-mechanisms","pwc_aliases":["attention-mechanisms-1"]}],"n_papers_tagged":20,"archive_num_papers":20,"papers_newest_first":[{"paper":null,"title":"To Bin or not to Bin: Alternative Representations of Mass Spectra","date":"2025-02-15","arxiv_id":"2502.10851","n_code_links":0,"syntology":null},{"paper":null,"title":"Advances in Set Function Learning: A Survey of Techniques and Applications","date":"2025-01-24","arxiv_id":"2501.14991","n_code_links":0,"syntology":null},{"paper":null,"title":"Adaptive parameters identification for nonlinear dynamics using deep permutation invariant networks","date":"2025-01-20","arxiv_id":"2501.11350","n_code_links":0,"syntology":null},{"paper":"/paper/multiset-transformer-advancing-representation","title":"Multiset Transformer: Advancing Representation Learning in Persistence Diagrams","date":"2024-11-22","arxiv_id":"2411.14662","n_code_links":1,"syntology":null},{"paper":"/paper/graph-as-point-set","title":"Graph as Point Set","date":"2024-05-05","arxiv_id":"2405.02795","n_code_links":1,"syntology":{"ran":20,"of":27,"unverified":7,"pointer_only":27}},{"paper":"/paper/associative-transformer-is-a-sparse","title":"Associative Transformer","date":"2023-09-22","arxiv_id":"2309.12862","n_code_links":1,"syntology":null},{"paper":null,"title":"Pointersect: Neural Rendering with Cloud-Ray Intersection","date":"2023-04-24","arxiv_id":"2304.12390","n_code_links":0,"syntology":null},{"paper":"/paper/event-voxel-set-transformer-for","title":"Event Voxel Set Transformer for Spatiotemporal Representation Learning on Event Streams","date":"2023-03-07","arxiv_id":"2303.03856","n_code_links":1,"syntology":null},{"paper":"/paper/set-norm-and-equivariant-skip-connections-1","title":"Set Norm and Equivariant Skip Connections: Putting the Deep in Deep Sets","date":"2022-06-23","arxiv_id":"2206.11925","n_code_links":1,"syntology":{"ran":1,"of":2,"unverified":1,"pointer_only":0}},{"paper":"/paper/permutation-invariant-relational-network-for","title":"Permutation-Invariant Relational Network for Multi-person 3D Pose Estimation","date":"2022-04-11","arxiv_id":"2204.04913","n_code_links":0,"syntology":null},{"paper":"/paper/voxel-set-transformer-a-set-to-set-approach","title":"Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds","date":"2022-03-19","arxiv_id":"2203.10314","n_code_links":1,"syntology":{"ran":0,"of":6,"unverified":6,"pointer_only":0}},{"paper":"/paper/automated-identification-of-cell-populations","title":"Automated Identification of Cell Populations in Flow Cytometry Data with Transformers","date":"2021-08-23","arxiv_id":"2108.10072","n_code_links":1,"syntology":null},{"paper":"/paper/you-are-allset-a-multiset-function-framework","title":"You are AllSet: A Multiset Function Framework for Hypergraph Neural Networks","date":"2021-06-24","arxiv_id":"2106.13264","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":"/paper/setvae-learning-hierarchical-composition-for","title":"SetVAE: Learning Hierarchical Composition for Generative Modeling of Set-Structured Data","date":"2021-03-29","arxiv_id":"2103.15619","n_code_links":2,"syntology":{"ran":17,"of":27,"unverified":10,"pointer_only":0}},{"paper":"/paper/latent-variable-nested-set-transformers","title":"Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction","date":"2021-02-19","arxiv_id":"2104.00563","n_code_links":2,"syntology":{"ran":2,"of":17,"unverified":15,"pointer_only":0}},{"paper":"/paper/gaining-insight-into-sars-cov-2-infection-and","title":"Gaining Insight into SARS-CoV-2 Infection and COVID-19 Severity Using Self-supervised Edge Features and Graph Neural Networks","date":"2020-06-23","arxiv_id":"2006.12971","n_code_links":1,"syntology":null},{"paper":"/paper/self-supervised-edge-features-for-improved","title":"Self-supervised edge features for improved Graph Neural Network training","date":"2020-06-23","arxiv_id":"2007.04777","n_code_links":1,"syntology":null},{"paper":"/paper/few-shot-learning-as-domain-adaptation","title":"Few-Shot Learning as Domain Adaptation: Algorithm and Analysis","date":"2020-02-06","arxiv_id":"2002.02050","n_code_links":0,"syntology":null},{"paper":"/paper/learning-set-equivariant-functions-with-swarm","title":"Learning Set-equivariant Functions with SWARM Mappings","date":"2019-06-22","arxiv_id":"1906.09400","n_code_links":1,"syntology":null},{"paper":"/paper/set-transformer-a-framework-for-attention","title":"Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks","date":"2018-10-01","arxiv_id":"1810.00825","n_code_links":9,"syntology":{"ran":5,"of":9,"unverified":4,"pointer_only":4}}],"papers_shown":20,"tasks":[{"task":"/task/node-classification","name":"Node Classification","papers":3},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/few-shot-image-classification","name":"Few-Shot Image Classification","papers":2},{"task":"/task/classification","name":"General Classification","papers":2},{"task":"/task/graph-attention","name":"Graph Attention","papers":2},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/representation-learning","name":"Representation Learning","papers":2},{"task":"/task/3d-multi-person-pose-estimation","name":"3D Multi-Person Pose Estimation","papers":1},{"task":"/task/3d-multi-person-pose-estimation-absolute","name":"3D Multi-Person Pose Estimation (absolute)","papers":1},{"task":"/task/3d-multi-person-pose-estimation-root-relative","name":"3D Multi-Person Pose Estimation (root-relative)","papers":1},{"task":"/task/3d-object-detection","name":"3D Object Detection","papers":1},{"task":"/task/3d-pose-estimation","name":"3D Pose Estimation","papers":1},{"task":"/task/3d-shape-recognition","name":"3D Shape Recognition","papers":1},{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":1},{"task":"/task/all","name":"All","papers":1},{"task":"/task/artificial-global-workspace","name":"Artificial Global Workspace","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/diversity","name":"Diversity","papers":1}],"tasks_shown":20,"n_tasks":48,"usage_by_year":[{"year":"2018","papers":1},{"year":"2019","papers":1},{"year":"2020","papers":3},{"year":"2021","papers":4},{"year":"2022","papers":3},{"year":"2023","papers":3},{"year":"2024","papers":2},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/set-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}