Papers › Objects do not disappear: Video object detection by single-frame object location anticipation

Objects do not disappear: Video object detection by single-frame object location anticipation

9 Aug 2023ICCV 2023 1arXiv:2308.04770archive 2025-07-28

Xin Liu, Fatemeh Karimi Nejadasl, Jan C. van Gemert, Olaf Booij, Silvia L. Pintea

Objects in videos are typically characterized by continuous smooth motion. We exploit continuous smooth motion in three ways. 1) Improved accuracy by using object motion as an additional source of supervision, which we obtain by anticipating object locations from a static keyframe. 2) Improved efficiency by only doing the expensive feature computations on a small subset of all frames. Because neighboring video frames are often redundant, we only compute features for a single static keyframe and predict object locations in subsequent frames. 3) Reduced annotation cost, where we only annotate the keyframe and use smooth pseudo-motion between keyframes. We demonstrate computational efficiency, annotation efficiency, and improved mean average precision compared to the state-of-the-art on four datasets: ImageNet VID, EPIC KITCHENS-55, YouTube-BoundingBoxes, and Waymo Open dataset. Our source code is available at https://github.com/L-KID/Videoobject-detection-by-location-anticipation.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

l-kid/video-object-detection-by-location-anticipation officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Computational EfficiencyObjectObject DetectionVideo Object Detectionobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Object Detection EPIC-KITCHENS-55 Ours (Faster RCNN) mAP@.5 41.7 #1 of 1 Archive leaderboard report
Video Object Detection ImageNet VID Ours (Def. DETR + SwinB) MAP 91.3 #3 of 33 Archive leaderboard report
Video Object Detection ImageNet VID Ours (Def. DETR + R101) MAP 87.9 #8 of 33 Archive leaderboard report
Video Object Detection ImageNet VID Ours (Faster RCNN + R101) MAP 87.2 #10 of 33 Archive leaderboard report
Video Object Detection Waymo Open Dataset AP 59.28 #1 of 1 Archive leaderboard report
Video Object Detection YT-BB mAP 59.8 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections