Papers › Cross-Drone Transformer Network for Robust Single Object Tracking

Cross-Drone Transformer Network for Robust Single Object Tracking

5 Jun 2023IEEE Transactions on Circuits and Systems for Video Technology 2023 6archive 2025-07-28

Guanlin Chen, Pengfei Zhu, Bing Cao, Xing Wang, QinGhua Hu

Drones have been widely used in a variety of applications, e.g., aerial photography and military security, because of their high maneuverability and broad views compared with fixed cameras. Multi-drone tracking systems can provide rich information about targets by collecting complementary video clips from different views, especially when targets are occluded or disappear in some views. However, it is challenging to handle cross-drone information interaction and multi-drone information fusion in multi-drone visual tracking. Recently, Transformer has shown significant advantages in automatically modeling the correlation between templates and search regions for visual tracking. To leverage its potential in multi-drone tracking, we propose a novel cross-drone Transformer network (TransMDOT) for visual object tracking tasks. The self-attention mechanism is used to automatically capture the correlation between multiple templates and the corresponding search region to achieve multi-drone feature fusion. During the tracking process, a cross-drone mapping mechanism is proposed by using the surrounding information of the drone with promising tracking status as reference, assisting drones that lost targets to re-calibrate, which implements real-time cross-drone information interaction. As the existing multi-drone evaluation metrics only consider spatial information while ignore temporal information, we further present a system perception index (SPFI) that combines both temporal and spatial information to evaluate the tracking status of multiple drones. Experiments on the MDOT dataset prove that TransMDOT greatly surpasses the state-of-the-art methods in both single-drone performance and multi-drone system fusion performance. Our code will be available on https://github.com/cgjacklin/transmdot.

PaperPDFCode

Code

cgjacklin/transmdot mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ObjectObject TrackingVisual Object TrackingVisual Tracking

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections