{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hybrid-lstm-and-encoder-decoder-architecture","title":"Hybrid LSTM and Encoder-Decoder Architecture for Detection of Image Forgeries","arxiv_id":"1903.02495","date":"2019-03-06","proceeding":null,"authors":["Jawadul H. Bappy","Cody Simons","Lakshmanan Nataraj","B. S. Manjunath","Amit K. Roy-Chowdhury"],"abstract":"With advanced image journaling tools, one can easily alter the semantic\nmeaning of an image by exploiting certain manipulation techniques such as\ncopy-clone, object splicing, and removal, which mislead the viewers. In\ncontrast, the identification of these manipulations becomes a very challenging\ntask as manipulated regions are not visually apparent. This paper proposes a\nhigh-confidence manipulation localization architecture which utilizes\nresampling features, Long-Short Term Memory (LSTM) cells, and encoder-decoder\nnetwork to segment out manipulated regions from non-manipulated ones.\nResampling features are used to capture artifacts like JPEG quality loss,\nupsampling, downsampling, rotation, and shearing. The proposed network exploits\nlarger receptive fields (spatial maps) and frequency domain correlation to\nanalyze the discriminative characteristics between manipulated and\nnon-manipulated regions by incorporating encoder and LSTM network. Finally,\ndecoder network learns the mapping from low-resolution feature maps to\npixel-wise predictions for image tamper localization. With predicted mask\nprovided by final layer (softmax) of the proposed architecture, end-to-end\ntraining is performed to learn the network parameters through back-propagation\nusing ground-truth masks. Furthermore, a large image splicing dataset is\nintroduced to guide the training process. The proposed method is capable of\nlocalizing image manipulations at pixel level with high precision, which is\ndemonstrated through rigorous experimentation on three diverse datasets.","url_abs":"http://arxiv.org/abs/1903.02495v1","url_pdf":"http://arxiv.org/pdf/1903.02495v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hybrid-lstm-and-encoder-decoder-architecture","repo_url":"https://github.com/jawadbappy/forgery_localization_HLED","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.02495","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}