Datasets › SVT

SVT (Street View Text Dataset)

archive 2025-07-28

The Street View Text (SVT) dataset was harvested from Google Street View. Image text in this data exhibits high variability and often has low resolution. In dealing with outdoor street level imagery, we note two characteristics. (1) Image text often comes from business signage and (2) business names are easily available through geographic business searches. These factors make the SVT set uniquely suited for word spotting in the wild: given a street view image, the goal is to identify words from nearby businesses.

Note: the dataset has undergone revision since the time it was evaluated in this publication. Please consult the ICCV2011 paper for most up-to-date results.

Source: Street View Text Dataset

Image source: Street View Text Dataset

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Scene Text Recognition SVT CLIP4STR-H (DFN-5B) Accuracy 99.1 CLIP4STR: A Simple Baseline for Scene Text Recognition... VamosC/CLIP4STR 37 Compare

Papers archive 2025-07-28

30 shown of 32 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 35. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
An Empirical Study of Scaling Law for OCR 1 1 29 Dec 2023 not harvested
DTrOCR: Decoder-only Transformer for Optical Character Recognition 1 1 30 Aug 2023 not harvested
Context Perception Parallel Decoder for Scene Text Recognition 2 1 23 Jul 2023 not harvested
DiffusionSTR: Diffusion Model for Scene Text Recognition 0 1 29 Jun 2023 not harvested
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model 1 4 23 May 2023 ran 3 of 10 samples (7 unverified)
TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text Recognition 1 1 9 May 2023 not harvested
Self-supervised Character-to-Character Distillation for Text Recognition 1 3 1 Nov 2022 not harvested
Multi-Granularity Prediction for Scene Text Recognition 3 1 8 Sep 2022 not harvested
Scene Text Recognition with Permuted Autoregressive Sequence Models 2 1 14 Jul 2022 ran 5 of 8 samples (3 unverified)
Self-supervised Implicit Glyph Attention for Text Recognition 1 1 7 Mar 2022 not harvested
SAFL: A Self-Attention Scene Text Recognizer with Focal Loss 1 1 1 Jan 2022 not harvested
Visual Semantics Allow for Textual Reasoning Better in Scene Text Recognition 1 1 24 Dec 2021 not harvested
Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features 3 1 30 Nov 2021 not harvested
CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition 3 1 22 Nov 2021 not harvested
Look Back Again: Dual Parallel Attention Network for Accurate and Robust Scene Text Recognition 2 1 1 Aug 2021 not harvested
Why You Should Try the Real Data for the Scene Text Recognition 1 1 29 Jul 2021 not harvested
Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition 1 1 13 Jun 2021 not harvested
Vision Transformer for Fast and Efficient Scene Text Recognition 3 1 18 May 2021 not harvested
Revisiting Classification Perspective on Scene Text Recognition 1 1 22 Feb 2021 not harvested
SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition 3 1 22 May 2020 not harvested
Towards Accurate Scene Text Recognition with Semantic Reasoning Networks 3 1 27 Mar 2020 not harvested
TextScanner: Reading Characters in Order for Robust Scene Text Recognition 0 1 28 Dec 2019 not harvested
Decoupled Attention Network for Text Recognition 5 1 21 Dec 2019 not harvested
On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention 2 1 10 Oct 2019 not harvested
What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis 13 1 3 Apr 2019 ran 0 of 19 samples (19 unverified)
Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition 8 1 2 Nov 2018 ran 0 of 8 samples (8 unverified)
Scene Text Recognition from Two-Dimensional Perspective 0 1 18 Sep 2018 not harvested
ASTER: An Attentional Scene Text Recognizer with Flexible Rectification 4 1 25 Jun 2018 not harvested
Star-net: A spatial attention residue network for scene text recognition. 1 1 20 Sep 2016 not harvested
Robust Scene Text Recognition with Automatic Rectification 5 1 12 Mar 2016 not harvested

The full list of 32 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • SVT

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections