Papers › When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset

When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset

14 Jul 2024arXiv:2407.10125archive 2025-07-28

Yi Zhang, Wang Zeng, Sheng Jin, Chen Qian, Ping Luo, Wentao Liu

Recent years have witnessed increasing research attention towards pedestrian detection by taking the advantages of different sensor modalities (e.g. RGB, IR, Depth, LiDAR and Event). However, designing a unified generalist model that can effectively process diverse sensor modalities remains a challenge. This paper introduces MMPedestron, a novel generalist model for multimodal perception. Unlike previous specialist models that only process one or a pair of specific modality inputs, MMPedestron is able to process multiple modal inputs and their dynamic combinations. The proposed approach comprises a unified encoder for modal representation and fusion and a general head for pedestrian detection. We introduce two extra learnable tokens, i.e. MAA and MAF, for adaptive multi-modal feature fusion. In addition, we construct the MMPD dataset, the first large-scale benchmark for multi-modal pedestrian detection. This benchmark incorporates existing public datasets and a newly collected dataset called EventPed, covering a wide range of sensor modalities including RGB, IR, Depth, LiDAR, and Event data. With multi-modal joint training, our model achieves state-of-the-art performance on a wide range of pedestrian detection benchmarks, surpassing leading models tailored for specific sensor modality. For example, it achieves 71.1 AP on COCO-Persons and 72.6 AP on LLVIP. Notably, our model achieves comparable performance to the InternImage-H model on CrowdHuman with 30x smaller parameters. Codes and data are available at https://github.com/BubblyYi/MMPedestron.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

BubblyYi/MMPedestron officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Object DetectionMultispectral Object DetectionObject DetectionPedestrian Detection

Datasets

Introduced by this paper, per the archive.

MMPD-Dataset

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multispectral Object Detection FLIR MMPedestron mAP50 86.4% #1 of 18 Archive leaderboard report
Object Detection CrowdHuman (full body) MMPedestron AP 97.1 #2 of 19 Archive leaderboard report
Object Detection CrowdHuman (full body) MMPedestron mMR 30.8 #2 of 19 Archive leaderboard report
Object Detection EventPed MMPedestron AP 79.0 #1 of 6 Archive leaderboard report
Object Detection InOutDoor MMPedestron AP 65.7 #1 of 6 Archive leaderboard report
Object Detection STCrowd MMPedestron AP 74.9 #1 of 6 Archive leaderboard report
Pedestrian Detection LLVIP MMPedestron AP 0.726 #1 of 15 Archive leaderboard report
Pedestrian Detection MMPD-Dataset MMPedestron box mAP 79.0 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections