{"url":"/method/pq-transformer","slug":"pq-transformer","name":"PQ-Transformer","full_name":"PointQuad-Transformer","full_name_withheld":false,"description_markdown":"**PQ-Transformer**, or **PointQuad-Transformer**, is a [Transformer](https://paperswithcode.com/method/transformer)-based architecture that predicts 3D objects and layouts simultaneously, using point cloud inputs. Unlike existing methods that either estimate layout keypoints or edges, room layouts are directly parameterized as a set of quads. Along with the quad representation, a physical constraint loss function is used that discourages object-layout interference.\r\n\r\nGiven an input 3D point cloud of $N$ points, the point cloud feature learning backbone extracts $M$ context-aware point features of $\\left(3+C\\right)$ dimensions, through sampling and grouping. A voting module and a farthest point sampling (FPS) module are used to generate $K\\_{1}$ object proposals and $K\\_{2}$ quad proposals respectively. Then the proposals are processed by a transformer decoder to further refine proposal features. Through several feedforward layers and non-maximum suppression (NMS), the proposals become the final object bounding boxes and layout quads.","description_state":"present","introduced_year":null,"introduced_by":{"title":"PQ-Transformer: Jointly Parsing 3D Objects and Layouts from Point Clouds","paper":"/paper/pq-transformer-jointly-parsing-3d-objects-and","first_author":"Xiaoxue Chen","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/pq-transformer-jointly-parsing-3d-objects-and"},"source":{"url":"https://arxiv.org/abs/2109.05566v2","title":"PQ-Transformer: Jointly Parsing 3D Objects and Layouts from Point Clouds","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Point Cloud Models","url":"/methods/category/point-cloud-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/pq-transformer-jointly-parsing-3d-objects-and","title":"PQ-Transformer: Jointly Parsing 3D Objects and Layouts from Point Clouds","date":"2021-09-12","arxiv_id":"2109.05566","n_code_links":1,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/room-layout-estimation","name":"Room Layout Estimation","papers":1},{"task":"/task/scene-understanding","name":"Scene Understanding","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/pq-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}