Papers › Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization

Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization

3 Sep 2024arXiv:2409.01726archive 2025-07-28

Qi Zhang, Kaiyi Zhang, Antoni B. Chan, Hui Huang

Multi-view crowd localization predicts the ground locations of all people in the scene. Typical methods usually estimate the crowd density maps on the ground plane first, and then obtain the crowd locations. However, the performance of existing methods is limited by the ambiguity of the density maps in crowded areas, where local peaks can be smoothed away. To mitigate the weakness of density map supervision, optimal transport-based point supervision methods have been proposed in the single-image crowd localization tasks, but have not been explored for multi-view crowd localization yet. Thus, in this paper, we propose a novel Mahalanobis distance-based multi-view optimal transport (M-MVOT) loss specifically designed for multi-view crowd localization. First, we replace the Euclidean-based transport cost with the Mahalanobis distance, which defines elliptical iso-contours in the cost function whose long-axis and short-axis directions are guided by the view ray direction. Second, the object-to-camera distance in each view is used to adjust the optimal transport cost of each location further, where the wrong predictions far away from the camera are more heavily penalized. Finally, we propose a strategy to consider all the input camera views in the model loss (M-MVOT) by computing the optimal transport cost for each ground-truth point based on its closest camera. Experiments demonstrate the advantage of the proposed method over density map-based or common Euclidean distance-based optimal transport loss on several multi-view crowd localization datasets. Project page: https://vcc.tech/research/2024/MVOT.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multiview Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multiview Detection CVCS M-MVOT MODA (0.5m) 43.5 #1 of 6 Archive leaderboard report
Multiview Detection CVCS M-MVOT MODA (1m) / #1 of 6 Archive leaderboard report
Multiview Detection MultiviewX M-MVOT MODA 96.7 #1 of 9 Archive leaderboard report
Multiview Detection MultiviewX M-MVOT MODP 86.1 #1 of 9 Archive leaderboard report
Multiview Detection MultiviewX M-MVOT Recall 97.9 #1 of 9 Archive leaderboard report
Multiview Detection Wildtrack M-MVOT MODA 92.1 #4 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections