Datasets › MeViS

MeViS (Motion expressions Video Segmentation)

Introduced by Henghui Ding et al. in MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions16 Aug 2023 archive 2025-07-28

MeViS is a large-scale dataset for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. The dataset contains numerous motion expressions to indicate target objects in complex environments.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Referring Video Object Segmentation MeViS MPG-SAM 2 J&F 53.7 MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global... rongfu-dsb/MPG-SAM2 16 Compare

Papers archive 2025-07-28

16 shown of 16 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 46. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation 1 1 10 Apr 2025 ran 3 of 12 samples (9 unverified; 12 pointer-only for licence)
Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation 1 1 5 Mar 2025 not harvested
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations 0 1 24 Jan 2025 not harvested
MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation 1 1 23 Jan 2025 ran 5 of 16 samples (11 unverified)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling 1 1 21 Jan 2025 not harvested
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation 1 1 15 Jan 2025 not harvested
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation 1 1 9 Jan 2025 not harvested
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation 1 1 26 Nov 2024 not harvested
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation 1 1 4 Apr 2024 ran 2 of 7 samples (5 unverified; 7 pointer-only for licence)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory 1 1 28 Mar 2024 ran 14 of 15 samples (1 unverified)
MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions 1 1 16 Aug 2023 ran 0 of 5 samples (5 unverified)
VLT: Vision-Language Transformer and Query Generation for Referring Segmentation 1 1 28 Oct 2022 ran 0 of 6 samples (6 unverified)
Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation 1 1 8 Jun 2022 ran 6 of 8 samples (2 unverified)
Language as Queries for Referring Video Object Segmentation 1 1 3 Jan 2022 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
End-to-End Referring Video Object Segmentation with Multimodal Transformers 2 1 29 Nov 2021 ran 6 of 11 samples (5 unverified)
URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark 1 1 1 Aug 2020 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MeViS

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections