Papers › Vision-Position Multi-Modal Beam Prediction Using Real Millimeter Wave Datasets

Vision-Position Multi-Modal Beam Prediction Using Real Millimeter Wave Datasets

15 Nov 2021arXiv:2111.07574archive 2025-07-28

Gouranga Charan, Tawfik Osman, Andrew Hredzak, Ngwe Thawdar, Ahmed Alkhateeb

Enabling highly-mobile millimeter wave (mmWave) and terahertz (THz) wireless communication applications requires overcoming the critical challenges associated with the large antenna arrays deployed at these systems. In particular, adjusting the narrow beams of these antenna arrays typically incurs high beam training overhead that scales with the number of antennas. To address these challenges, this paper proposes a multi-modal machine learning based approach that leverages positional and visual (camera) data collected from the wireless communication environment for fast beam prediction. The developed framework has been tested on a real-world vehicular dataset comprising practical GPS, camera, and mmWave beam training data. The results show the proposed approach achieves more than ≈ 75\% top-1 beam prediction accuracy and close to 100\% top-3 beam prediction accuracy in realistic communication scenarios.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Beam PredictionPrediction

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

GPS

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections