Papers › Adaptive Aspect Ratios with Patch-Mixup-ViT-based Vehicle ReID

Adaptive Aspect Ratios with Patch-Mixup-ViT-based Vehicle ReID

9 Nov 2024arXiv:2411.06297archive 2025-07-28

Mei Qiu, Lauren Ann Christopher, Stanley Chien, Lingxi Li

Vision Transformers (ViTs) have shown exceptional performance in vehicle re-identification (ReID) tasks. However, non-square aspect ratios of image or video inputs can negatively impact re-identification accuracy. To address this challenge, we propose a novel, human perception driven, and general ViT-based ReID framework that fuses models trained on various aspect ratios. Our key contributions are threefold: (i) We analyze the impact of aspect ratios on performance using the VeRi-776 and VehicleID datasets, providing guidance for input settings based on the distribution of original image aspect ratios. (ii) We introduce patch-wise mixup strategy during ViT patchification (guided by spatial attention scores) and implement uneven stride for better alignment with object aspect ratios. (iii) We propose a dynamic feature fusion ReID network to enhance model robustness. Our method outperforms state-of-the-art transformer-based approaches on both datasets, with only a minimal increase in inference time per image.

PaperPDFCode

Code

qiumei1101/adaptive_ar_pm_transreid officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Vehicle Re-Identification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AttentionMixupSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections