{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/t-cnn-tubelets-with-convolutional-neural","title":"T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos","arxiv_id":"1604.02532","date":"2016-04-09","proceeding":null,"authors":["Kai Kang","Hongsheng Li","Junjie Yan","Xingyu Zeng","Bin Yang","Tong Xiao","Cong Zhang","Zhe Wang","Ruohui Wang","Xiaogang Wang","Wanli Ouyang"],"abstract":"The state-of-the-art performance for object detection has been significantly\nimproved over the past two years. Besides the introduction of powerful deep\nneural networks such as GoogleNet and VGG, novel object detection frameworks\nsuch as R-CNN and its successors, Fast R-CNN and Faster R-CNN, play an\nessential role in improving the state-of-the-art. Despite their effectiveness\non still images, those frameworks are not specifically designed for object\ndetection from videos. Temporal and contextual information of videos are not\nfully investigated and utilized. In this work, we propose a deep learning\nframework that incorporates temporal and contextual information from tubelets\nobtained in videos, which dramatically improves the baseline performance of\nexisting still-image detection frameworks when they are applied to videos. It\nis called T-CNN, i.e. tubelets with convolutional neueral networks. The\nproposed framework won the recently introduced object-detection-from-video\n(VID) task with provided data in the ImageNet Large-Scale Visual Recognition\nChallenge 2015 (ILSVRC2015).","url_abs":"http://arxiv.org/abs/1604.02532v4","url_pdf":"http://arxiv.org/pdf/1604.02532v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"t-cnn-tubelets-with-convolutional-neural","repo_url":"https://github.com/myfavouritekk/T-CNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"novel-object-detection","task_name":"Novel Object Detection"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"auxiliary-classifier","method_name":"Auxiliary Classifier"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"fast-r-cnn","method_name":"Fast R-CNN"},{"method_slug":"faster-r-cnn","method_name":"Faster R-CNN"},{"method_slug":"googlenet","method_name":"GoogLeNet"},{"method_slug":"inception-module","method_name":"Inception Module"},{"method_slug":"local-response-normalization","method_name":"Local Response Normalization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"roipool","method_name":"RoIPool"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.02532","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1604.02532"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/myfavouritekk/T-CNN","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f9c5f5b57e852959","entry":"proto_load","repo":"myfavouritekk/T-CNN","repo_kind":"official","path":"track_data_layer/layer.py","file_url":"https://github.com/myfavouritekk/T-CNN/blob/HEAD/track_data_layer/layer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f9c5f5b57e852959"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}