Papers › Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation)

Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation)

12 Aug 2021arXiv:2108.05888archive 2025-07-28

Yunzhong Hou, Liang Zheng

Multiview detection incorporates multiple camera views to deal with occlusions, and its central problem is multiview aggregation. Given feature map projections from multiple views onto a common ground plane, the state-of-the-art method addresses this problem via convolution, which applies the same calculation regardless of object locations. However, such translation-invariant behaviors might not be the best choice, as object features undergo various projection distortions according to their positions and cameras. In this paper, we propose a novel multiview detector, MVDeTr, that adopts a newly introduced shadow transformer to aggregate multiview information. Unlike convolutions, shadow transformer attends differently at different positions and cameras to deal with various shadow-like distortions. We propose an effective training scheme that includes a new view-coherent data augmentation method, which applies random augmentations while maintaining multiview consistency. On two multiview detection benchmarks, we report new state-of-the-art accuracy with the proposed system. Code is available at https://github.com/hou-yz/MVDeTr.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

hou-yz/mvdetr officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Data AugmentationMultiview DetectionTranslation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multiview Detection CVCS MVDeTr F1_score (1m) 61.0 #4 of 6 Archive leaderboard report
Multiview Detection CVCS MVDeTr MODA (1m) 39.8 #4 of 6 Archive leaderboard report
Multiview Detection CVCS MVDeTr MODP (1m) 84.1 #4 of 6 Archive leaderboard report
Multiview Detection CVCS MVDeTr Precision (1m) 95.3 #4 of 6 Archive leaderboard report
Multiview Detection CVCS MVDeTr Recall (1m) 44.9 #4 of 6 Archive leaderboard report
Multiview Detection CityStreet MVDeTr F1_score (2m) 75.2 #2 of 5 Archive leaderboard report
Multiview Detection CityStreet MVDeTr MODA (2m) 58.3 #2 of 5 Archive leaderboard report
Multiview Detection CityStreet MVDeTr MODP (2m) 74.1 #2 of 5 Archive leaderboard report
Multiview Detection CityStreet MVDeTr Precision (2m) 92.8 #2 of 5 Archive leaderboard report
Multiview Detection CityStreet MVDeTr Recall (2m) 63.2 #2 of 5 Archive leaderboard report
Multiview Detection MultiviewX MVDeTr MODA 93.7 #5 of 9 Archive leaderboard report
Multiview Detection MultiviewX MVDeTr MODP 91.3 #5 of 9 Archive leaderboard report
Multiview Detection MultiviewX MVDeTr Recall 94.2 #5 of 9 Archive leaderboard report
Multiview Detection Wildtrack MVDeTr MODA 91.5 #6 of 10 Archive leaderboard report
Multiview Detection Wildtrack MVDeTr MODP 82.1 #6 of 10 Archive leaderboard report
Multiview Detection Wildtrack MVDeTr Recall 94.0 #6 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections