Papers › GLAMI-1M: A Multilingual Image-Text Fashion Dataset

GLAMI-1M: A Multilingual Image-Text Fashion Dataset

17 Nov 2022BMVC 2022 11arXiv:2211.14451archive 2025-07-28

Vaclav Kosar, Antonín Hoskovec, Milan Šulc, Radek Bartyzal

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text. The dataset, source code and model checkpoints are published at https://github.com/glami/glami-1m

PaperPDFConference PDFCode

Code

glami/glami-1m officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationImage GenerationImage-text ClassificationMultilingual Image-Text ClassificationText Classificationtext-classification

Datasets

Introduced by this paper, per the archive.

GLAMI-1M

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multilingual Image-Text Classification GLAMI-1M EmbraceNet (image+text) Top 1 Accuracy % 69.7 #1 of 2 Archive leaderboard report
Multilingual Image-Text Classification GLAMI-1M EmbraceNet (image+text) Top 5 Accuracy % 94.0 #1 of 2 Archive leaderboard report
Multilingual Image-Text Classification GLAMI-1M CLIP (zero-shot image+text) Top 1 Accuracy % 32.3 #2 of 2 Archive leaderboard report
Multilingual Image-Text Classification GLAMI-1M CLIP (zero-shot image+text) Top 5 Accuracy % 74.5 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAdafactorAttentionAttention DropoutAverage PoolingBPEBatch NormalizationBottleneck Residual BlockCLIPConvolutionDense ConnectionsDiffusionDropoutEmbraceNetGated Linear UnitGlobal Average PoolingInverse Square Root ScheduleKaiming InitializationLayer NormalizationLinear LayerMax PoolingMulti-Head AttentionReLUResidual BlockResidual ConnectionSentencePieceSoftmaxT5TestmT5

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections