Papers › Towards Streaming Perception

Towards Streaming Perception

21 May 2020ECCV 2020 8arXiv:2005.10420archive 2025-07-28

Mengtian Li, Yu-Xiong Wang, Deva Ramanan

Embodied perception refers to the ability of an autonomous agent to perceive its environment so that it can (re)act. The responsiveness of the agent is largely governed by latency of its processing pipeline. While past work has studied the algorithmic trade-off between latency and accuracy, there has not been a clear metric to compare different methods along the Pareto optimal latency-accuracy curve. We point out a discrepancy between standard offline evaluation and real-time applications: by the time an algorithm finishes processing a particular frame, the surrounding world has changed. To these ends, we present an approach that coherently integrates latency and accuracy into a single metric for real-time online perception, which we refer to as "streaming accuracy". The key insight behind this metric is to jointly evaluate the output of the entire perception stack at every time instant, forcing the stack to consider the amount of streaming data that should be ignored while computation is occurring. More broadly, building upon this metric, we introduce a meta-benchmark that systematically converts any single-frame task into a streaming perception task. We focus on the illustrative tasks of object detection and instance segmentation in urban video streams, and contribute a novel dataset with high-quality and temporally-dense annotations. Our proposed solutions and their empirical analysis demonstrate a number of surprising conclusions: (1) there exists an optimal "sweet spot" that maximizes streaming accuracy along the Pareto optimal latency-accuracy curve, (2) asynchronous tracking and future forecasting naturally emerge as internal representations that enable streaming perception, and (3) dynamic scheduling can be used to overcome temporal aliasing, yielding the paradoxical result that latency is sometimes minimized by sitting idle and "doing nothing".

PaperPDFConference PDFCode

Code

mtli/sAP mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Instance SegmentationMotion ForecastingObject DetectionReal-Time Multi-Object TrackingReal-Time Object DetectionSchedulingSemantic Segmentationobject-detection

Datasets

Introduced by this paper, per the archive.

Argoverse-HD

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Real-Time Object Detection Argoverse-HD (Detection-Only, Test) Official challenge baseline AP 13.61 #3 of 3 Archive leaderboard report
Real-Time Object Detection Argoverse-HD (Detection-Only, Val) Official challenge baseline AP 14.91 #2 of 2 Archive leaderboard report
Real-Time Object Detection Argoverse-HD (Full-Stack, Test) Official challenge baseline AP 21.06 #3 of 3 Archive leaderboard report
Real-Time Object Detection Argoverse-HD (Full-Stack, Val) Official challenge baseline AP 21.06 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections