{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/object-detection-from-video-tubelets-with","title":"Object Detection from Video Tubelets with Convolutional Neural Networks","arxiv_id":"1604.04053","date":"2016-04-14","proceeding":"CVPR 2016 6","authors":["Kai Kang","Wanli Ouyang","Hongsheng Li","Xiaogang Wang"],"abstract":"Deep Convolution Neural Networks (CNNs) have shown impressive performance in\nvarious vision tasks such as image classification, object detection and\nsemantic segmentation. For object detection, particularly in still images, the\nperformance has been significantly increased last year thanks to powerful deep\nnetworks (e.g. GoogleNet) and detection frameworks (e.g. Regions with CNN\nfeatures (R-CNN)). The lately introduced ImageNet task on object detection from\nvideo (VID) brings the object detection task into the video domain, in which\nobjects' locations at each frame are required to be annotated with bounding\nboxes. In this work, we introduce a complete framework for the VID task based\non still-image object detection and general object tracking. Their relations\nand contributions in the VID task are thoroughly studied and evaluated. In\naddition, a temporal convolution network is proposed to incorporate temporal\ninformation to regularize the detection results and shows its effectiveness for\nthe task.","url_abs":"http://arxiv.org/abs/1604.04053v1","url_pdf":"http://arxiv.org/pdf/1604.04053v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"object-detection-from-video-tubelets-with","repo_url":"https://github.com/myfavouritekk/vdetlib","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.04053","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}