{"url":"/method/tum","slug":"tum","name":"TUM","full_name":"Thinned U-shape Module","full_name_withheld":false,"description_markdown":"**Thinned U-shape Module**, or **TUM**, is a feature extraction block used for object detection models. It was introduced as part of the [M2Det](https://paperswithcode.com/method/m2det) architecture. Different from [FPN](https://paperswithcode.com/method/fpn) and [RetinaNet](https://paperswithcode.com/method/retinanet), TUM adopts a thinner U-shape structure as illustrated in the Figure to the right. The encoder is a series of 3x3 [convolution](https://paperswithcode.com/method/convolution) layers with stride 2. And the decoder takes the outputs of these layers as its reference set of feature maps, while the original FPN chooses the output of the last layer of each stage in [ResNet](https://paperswithcode.com/method/resnet) backbone. \r\n\r\nIn addition, with TUM, we add [1x1 convolution](https://paperswithcode.com/method/1x1-convolution) layers after the upsample and element-wise sum operation at the decoder branch to enhance learning ability and keep smoothness for the features. In the context of M2Det, all of the outputs in the decoder of each TUM form the multi-scale features of the current level. As a whole, the outputs of stacked TUMs form the multi-level multi-scale features, while the front TUM mainly provides shallow-level features, the middle TUM provides medium-level features, and the back TUM provides deep-level features.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1811.04533v3","title":"M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/qijiezhao/M2Det/blob/ade4f3d12979800c367bf1e46d2e316e73a87514/layers/nn_utils.py#L26","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Feature Extractors","url":"/methods/category/feature-extractors","pwc_aliases":[]}],"n_papers_tagged":57,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"GS4: Generalizable Sparse Splatting Semantic SLAM","date":"2025-06-06","arxiv_id":"2506.06517","n_code_links":0,"syntology":null},{"paper":null,"title":"Black-box Adversarial Attacks on CNN-based SLAM Algorithms","date":"2025-05-30","arxiv_id":"2505.24654","n_code_links":0,"syntology":null},{"paper":null,"title":"FGS-SLAM: Fourier-based Gaussian Splatting for Real-time SLAM with Sparse and Dense Map Fusion","date":"2025-03-03","arxiv_id":"2503.01109","n_code_links":0,"syntology":null},{"paper":"/paper/surface-sos-self-supervised-object","title":"Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation","date":"2025-01-17","arxiv_id":"2501.09947","n_code_links":1,"syntology":null},{"paper":null,"title":"AutoLoop: Fast Visual SLAM Fine-tuning through Agentic Curriculum Learning","date":"2025-01-15","arxiv_id":"2501.09160","n_code_links":0,"syntology":null},{"paper":null,"title":"SP-SLAM: Neural Real-Time Dense SLAM With Scene Priors","date":"2025-01-11","arxiv_id":"2501.06469","n_code_links":0,"syntology":null},{"paper":null,"title":"Scaffold-SLAM: Structured 3D Gaussians for Simultaneous Localization and Photorealistic Mapping","date":"2025-01-09","arxiv_id":"2501.05242","n_code_links":0,"syntology":null},{"paper":"/paper/gsplatloc-ultra-precise-camera-localization","title":"GSplatLoc: Ultra-Precise Camera Localization via 3D Gaussian Splatting","date":"2024-12-28","arxiv_id":"2412.20056","n_code_links":1,"syntology":null},{"paper":null,"title":"LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models","date":"2024-12-10","arxiv_id":"2412.07746","n_code_links":0,"syntology":null},{"paper":null,"title":"A Semantic Communication System for Real-time 3D Reconstruction Tasks","date":"2024-12-02","arxiv_id":"2412.01191","n_code_links":0,"syntology":null},{"paper":null,"title":"DATAP-SfM: Dynamic-Aware Tracking Any Point for Robust Structure from Motion in the Wild","date":"2024-11-20","arxiv_id":"2411.13291","n_code_links":0,"syntology":null},{"paper":"/paper/v3d-slam-robust-rgb-d-slam-in-dynamic","title":"V3D-SLAM: Robust RGB-D SLAM in Dynamic Environments with 3D Semantic Geometry Voting","date":"2024-10-15","arxiv_id":"2410.12068","n_code_links":1,"syntology":null},{"paper":null,"title":"CLIP-Clique: Graph-based Correspondence Matching Augmented by Vision Language Models for Object-based Global Localization","date":"2024-10-04","arxiv_id":"2410.03054","n_code_links":0,"syntology":null},{"paper":"/paper/inline-photometrically-calibrated-hybrid","title":"Inline Photometrically Calibrated Hybrid Visual SLAM","date":"2024-09-25","arxiv_id":"2409.16810","n_code_links":1,"syntology":null},{"paper":"/paper/towards-global-localization-using-multi-modal","title":"Towards Global Localization using Multi-Modal Object-Instance Re-Identification","date":"2024-09-18","arxiv_id":"2409.12002","n_code_links":1,"syntology":null},{"paper":null,"title":"UDGS-SLAM : UniDepth Assisted Gaussian Splatting for Monocular SLAM","date":"2024-08-31","arxiv_id":"2409.00362","n_code_links":0,"syntology":null},{"paper":null,"title":"Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction","date":"2024-08-28","arxiv_id":"2408.15608","n_code_links":0,"syntology":null},{"paper":null,"title":"Transfer Learning from Simulated to Real Scenes for Monocular 3D Object Detection","date":"2024-08-28","arxiv_id":"2408.15637","n_code_links":0,"syntology":null},{"paper":null,"title":"FAWN: Floor-And-Walls Normal Regularization for Direct Neural TSDF Reconstruction","date":"2024-06-17","arxiv_id":"2406.12054","n_code_links":0,"syntology":null},{"paper":null,"title":"MEDeA: Multi-view Efficient Depth Adjustment","date":"2024-06-17","arxiv_id":"2406.12048","n_code_links":0,"syntology":null},{"paper":null,"title":"Predictive Energy Management for Battery Electric Vehicles with Hybrid Models","date":"2024-05-15","arxiv_id":"2405.10984","n_code_links":0,"syntology":null},{"paper":null,"title":"SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping","date":"2024-04-17","arxiv_id":"2404.11419","n_code_links":0,"syntology":null},{"paper":"/paper/high-fidelity-slam-using-gaussian-splatting","title":"High-Fidelity SLAM Using Gaussian Splatting with Rendering-Guided Densification and Regularized Optimization","date":"2024-03-19","arxiv_id":"2403.12535","n_code_links":1,"syntology":null},{"paper":null,"title":"An Error-Matching Exclusion Method for Accelerating Visual SLAM","date":"2024-02-22","arxiv_id":"2402.14345","n_code_links":0,"syntology":null},{"paper":null,"title":"A Robust Error-Resistant View Selection Method for 3D Reconstruction","date":"2024-02-18","arxiv_id":"2402.11431","n_code_links":0,"syntology":null},{"paper":"/paper/activeanno3d-an-active-learning-framework-for","title":"ActiveAnno3D -- An Active Learning Framework for Multi-Modal 3D Object Detection","date":"2024-02-05","arxiv_id":"2402.03235","n_code_links":1,"syntology":null},{"paper":null,"title":"Geometry Depth Consistency in RGBD Relative Pose Estimation","date":"2024-01-01","arxiv_id":"2401.00639","n_code_links":0,"syntology":null},{"paper":null,"title":"3DS-SLAM: A 3D Object Detection based Semantic SLAM towards Dynamic Indoor Environments","date":"2023-10-10","arxiv_id":"2310.06385","n_code_links":0,"syntology":null},{"paper":"/paper/dynamon-motion-aware-fast-and-robust-camera","title":"DynaMoN: Motion-Aware Fast and Robust Camera Localization for Dynamic Neural Radiance Fields","date":"2023-09-16","arxiv_id":"2309.08927","n_code_links":1,"syntology":null},{"paper":"/paper/h-slam-hybrid-direct-indirect-visual-slam","title":"H-SLAM: Hybrid Direct-Indirect Visual SLAM","date":"2023-06-12","arxiv_id":"2306.07363","n_code_links":2,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/simultaneous-localization-and-mapping","name":"Simultaneous Localization and Mapping","papers":12},{"task":"/task/pose-estimation","name":"Pose Estimation","papers":10},{"task":"/task/3d-reconstruction","name":"3D Reconstruction","papers":7},{"task":"/task/object-detection","name":"Object Detection","papers":7},{"task":"/task/object-detection-1","name":"object-detection","papers":7},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":6},{"task":"/task/camera-localization","name":"Camera Localization","papers":5},{"task":"/task/object","name":"Object","papers":5},{"task":"/task/depth-estimation","name":"Depth Estimation","papers":4},{"task":"/task/nerf","name":"NeRF","papers":4},{"task":"/task/optical-flow-estimation","name":"Optical Flow Estimation","papers":4},{"task":"/task/visual-odometry","name":"Visual Odometry","papers":4},{"task":"/task/3d-object-detection","name":"3D Object Detection","papers":3},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":3},{"task":null,"name":"GPU","papers":3},{"task":"/task/semantic-slam","name":"Semantic SLAM","papers":3},{"task":"/task/surface-reconstruction","name":"Surface Reconstruction","papers":3},{"task":"/task/camera-pose-estimation","name":"Camera Pose Estimation","papers":2},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/novel-view-synthesis","name":"Novel View Synthesis","papers":2}],"tasks_shown":20,"n_tasks":61,"usage_by_year":[{"year":"2018","papers":1},{"year":"2020","papers":6},{"year":"2021","papers":6},{"year":"2022","papers":12},{"year":"2023","papers":5},{"year":"2024","papers":20},{"year":"2025","papers":7}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tum"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}