{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/integrated-object-detection-and-tracking-with","title":"Integrated Object Detection and Tracking with Tracklet-Conditioned Detection","arxiv_id":"1811.11167","date":"2018-11-27","proceeding":null,"authors":["Zheng Zhang","Dazhi Cheng","Xizhou Zhu","Stephen Lin","Jifeng Dai"],"abstract":"Accurate detection and tracking of objects is vital for effective video\nunderstanding. In previous work, the two tasks have been combined in a way that\ntracking is based heavily on detection, but the detection benefits marginally\nfrom the tracking. To increase synergy, we propose to more tightly integrate\nthe tasks by conditioning the object detection in the current frame on\ntracklets computed in prior frames. With this approach, the object detection\nresults not only have high detection responses, but also improved coherence\nwith the existing tracklets. This greater coherence leads to estimated object\ntrajectories that are smoother and more stable than the jittered paths obtained\nwithout tracklet-conditioned detection. Over extensive experiments, this\napproach is shown to achieve state-of-the-art performance in terms of both\ndetection and tracking accuracy, as well as noticeable improvements in tracking\nstability.","url_abs":"http://arxiv.org/abs/1811.11167v1","url_pdf":"http://arxiv.org/pdf/1811.11167v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"video-object-detection","task_name":"Video Object Detection"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-object-detection-on-imagenet-vid","task":"Video Object Detection","dataset":"ImageNet VID","model":"Tracklet-Conditioned Detection+DCNv2+FGFA","rank_in_archive_order":22,"of":33,"metrics":{"MAP ":"83.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.11167","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}