{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lstm-cf-unifying-context-modeling-and-fusion","title":"LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling","arxiv_id":"1604.05000","date":"2016-04-18","proceeding":null,"authors":["Zhen Li","Yukang Gan","Xiaodan Liang","Yizhou Yu","Hui Cheng","Liang Lin"],"abstract":"Semantic labeling of RGB-D scenes is crucial to many intelligent applications\nincluding perceptual robotics. It generates pixelwise and fine-grained label\nmaps from simultaneously sensed photometric (RGB) and depth channels. This\npaper addresses this problem by i) developing a novel Long Short-Term Memorized\nContext Fusion (LSTM-CF) Model that captures and fuses contextual information\nfrom multiple channels of photometric and depth data, and ii) incorporating\nthis model into deep convolutional neural networks (CNNs) for end-to-end\ntraining. Specifically, contexts in photometric and depth channels are,\nrespectively, captured by stacking several convolutional layers and a long\nshort-term memory layer; the memory layer encodes both short-range and\nlong-range spatial dependencies in an image along the vertical direction.\nAnother long short-term memorized fusion layer is set up to integrate the\ncontexts along the vertical direction from different channels, and perform\nbi-directional propagation of the fused vertical contexts along the horizontal\ndirection to obtain true 2D global contexts. At last, the fused contextual\nrepresentation is concatenated with the convolutional features extracted from\nthe photometric channels in order to improve the accuracy of fine-scale\nsemantic labeling. Our proposed model has set a new state of the art, i.e.,\n48.1% and 49.4% average class accuracy over 37 categories (2.2% and 5.4%\nimprovement) on the large-scale SUNRGBD dataset and the NYUDv2dataset,\nrespectively.","url_abs":"http://arxiv.org/abs/1604.05000v3","url_pdf":"http://arxiv.org/pdf/1604.05000v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lstm-cf-unifying-context-modeling-and-fusion","repo_url":"https://github.com/icemansina/LSTM-CF","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"scene-labeling","task_name":"Scene Labeling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}