Papers › Visual Keyword Spotting with Attention

Visual Keyword Spotting with Attention

29 Oct 2021arXiv:2110.15957archive 2025-07-28

K R Prajwal, Liliane Momeni, Triantafyllos Afouras, Andrew Zisserman

In this paper, we consider the task of spotting spoken keywords in silent video sequences -- also known as visual keyword spotting. To this end, we investigate Transformer-based models that ingest two streams, a visual encoding of the video and a phonetic encoding of the keyword, and output the temporal location of the keyword if present. Our contributions are as follows: (1) We propose a novel architecture, the Transpotter, that uses full cross-modal attention between the visual and phonetic streams; (2) We show through extensive evaluations that our model outperforms the prior state-of-the-art visual keyword spotting and lip reading methods on the challenging LRW, LRS2, LRS3 datasets by a large margin; (3) We demonstrate the ability of our model to spot words under the extreme conditions of isolated mouthings in sign language videos.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

prajwalkr/transpotter officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Lip ReadingVisual Keyword Spotting

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Keyword Spotting LRS2 Transpotter Top-1 Accuracy 65 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS2 Transpotter Top-5 Accuracy 87.1 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS2 Transpotter mAP 69.2 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS2 Transpotter mAP IOU@0.5 68.3 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS3-TED Transpotter Top-1 Accuracy 52 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS3-TED Transpotter Top-5 Accuracy 77.1 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS3-TED Transpotter mAP 55.4 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRS3-TED Transpotter mAP IOU@0.5 53.6 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRW Transpotter Top-1 Accuracy 85.8 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRW Transpotter Top-5 Accuracy 99.6 #1 of 1 Archive leaderboard report
Visual Keyword Spotting LRW Transpotter mAP 64.1 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections