Datasets › RefMatte

RefMatte (Referring Image Matting)

Introduced by Jizhizi Li et al. in Referring Image Matting10 Jun 2022 archive 2025-07-28

RefMatte is the first large-scale challenging dataset under the task referring image matting, generated by a comprehensive image composition and expression generation engine on top of current public high-quality matting foregrounds with flexible logics and re-labelled diverse attributes. RefMatte consists of 230 object categories, 47,500 images, 118,749 expression-region entities, and 474,996 expressions, which can be further extended easily in the future.

RefMatte comes along with two settings: keyword-based and expression-based. The former one takes a high-resolution image and a keyword as input, while the latter one takes a high-resolution image and a flowery expression as input.

Additionally, we construct a real-world test set with 100 high-resolution natural images and manually annotate complex phrases to evaluate the out-of-domain generalization abilities of RIM methods, named as RefMatte-RW100.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Referring Image Matting (Expression-based) RefMatte CLIPMat (ViT-L/14) SAD 42.05 Referring Image Matting jizhizili/rim 4 Compare
Referring Image Matting (Keyword-based) RefMatte CLIPMat (ViT-L/14) SAD 8.51 Referring Image Matting jizhizili/rim 4 Compare
Referring Image Matting (RefMatte-RW100) RefMatte CLIPMat (ViT-L/14) SAD 88.52 Referring Image Matting jizhizili/rim 4 Compare

Papers archive 2025-07-28

3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 7. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Referring Image Matting 1 6 10 Jun 2022 not harvested
Image Segmentation Using Text and Image Prompts 6 3 18 Dec 2021 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 5 3 26 Apr 2021 ran 6 of 11 samples (5 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • RefMatte

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections