{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-adversarial-visual-level-domain","title":"Unsupervised Adversarial Visual Level Domain Adaptation for Learning Video Object Detectors from Images","arxiv_id":"1810.02074","date":"2018-10-04","proceeding":null,"authors":["Avisek Lahiri","Charan Reddy","Prabir Kumar Biswas"],"abstract":"Deep learning based object detectors require thousands of diversified\nbounding box and class annotated examples. Though image object detectors have\nshown rapid progress in recent years with the release of multiple large-scale\nstatic image datasets, object detection on videos still remains an open problem\ndue to scarcity of annotated video frames. Having a robust video object\ndetector is an essential component for video understanding and curating\nlarge-scale automated annotations in videos. Domain difference between images\nand videos makes the transferability of image object detectors to videos\nsub-optimal. The most common solution is to use weakly supervised annotations\nwhere a video frame has to be tagged for presence/absence of object categories.\nThis still takes up manual effort. In this paper we take a step forward by\nadapting the concept of unsupervised adversarial image-to-image translation to\nperturb static high quality images to be visually indistinguishable from a set\nof video frames. We assume the presence of a fully annotated static image\ndataset and an unannotated video dataset. Object detector is trained on\nadversarially transformed image dataset using the annotations of the original\ndataset. Experiments on Youtube-Objects and Youtube-Objects-Subset datasets\nwith two contemporary baseline object detectors reveal that such unsupervised\npixel level domain adaptation boosts the generalization performance on video\nframes compared to direct application of original image object detector. Also,\nwe achieve competitive performance compared to recent baselines of weakly\nsupervised methods. This paper can be seen as an application of image\ntranslation for cross domain object detection.","url_abs":"http://arxiv.org/abs/1810.02074v1","url_pdf":"http://arxiv.org/pdf/1810.02074v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-adversarial-visual-level-domain","repo_url":"https://github.com/avisekiit/wacv_2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}