{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporally-coherent-video-harmonization-using","title":"Temporally Coherent Video Harmonization Using Adversarial Networks","arxiv_id":"1809.01372","date":"2018-09-05","proceeding":null,"authors":["Hao-Zhi Huang","Senzhe Xu","Junxiong Cai","Wei Liu","Shi-Min Hu"],"abstract":"Compositing is one of the most important editing operations for images and\nvideos. The process of improving the realism of composite results is often\ncalled harmonization. Previous approaches for harmonization mainly focus on\nimages. In this work, we take one step further to attack the problem of video\nharmonization. Specifically, we train a convolutional neural network in an\nadversarial way, exploiting a pixel-wise disharmony discriminator to achieve\nmore realistic harmonized results and introducing a temporal loss to increase\ntemporal consistency between consecutive harmonized frames. Thanks to the\npixel-wise disharmony discriminator, we are also able to relieve the need of\ninput foreground masks. Since existing video datasets which have ground-truth\nforeground masks and optical flows are not sufficiently large, we propose a\nsimple yet efficient method to build up a synthetic dataset supporting\nsupervised training of the proposed adversarial network. Experiments show that\ntraining on our synthetic dataset generalizes well to the real-world composite\ndataset. Also, our method successfully incorporates temporal consistency during\ntraining and achieves more harmonious results than previous methods.","url_abs":"http://arxiv.org/abs/1809.01372v1","url_pdf":"http://arxiv.org/pdf/1809.01372v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporally-coherent-video-harmonization-using","repo_url":"https://github.com/bcmi/video-harmonization-dataset-hyoutube","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"video-harmonization","task_name":"Video Harmonization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.01372","atlas_url":"https://app.syntology.ai/?focus=1809.01372","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}