{"url":"/dataset/m-3-vos","name":"M$^3$-VOS","full_name":"M$^3$-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation","description_markdown":"## 💡 Description\r\n\r\nA new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M$^3$-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10 distinct everyday scenarios. We collected 205,181 masks, with an average track duration of 14.27s. M$^3$-VOS covers 120+ categories of objects across 6 phases within 14 scenarios, encompassing 23 specific phase transitions. \r\n\r\n- **Venue:** CVPR2025\r\n- **Repository:** [Tool 🛠️](https://github.com/Lijiaxin0111/SemiAuto-Multi-Level-Annotation-Tool), [Page🏠](https://zixuan-chen.github.io/M-cube-VOS.github.io/)\r\n- **Paper:** arxiv.org/html/2412.13803v2\r\n- **Point of Contact:** [Jiaxin Li](li_jiaxin@sjtu.edu.cn) , [Zixuan Chen](m13953842591@sjtu.edu.cn)","description_withheld":null,"homepage":"https://zixuan-chen.github.io/M-cube-VOS.github.io/","introduced_date":"2025-06-15","introduced_date_note":null,"introduced_by":null,"license":{"name":"apache-2.0","url":"https://www.apache.org/licenses/LICENSE-2.0"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Video Object Segmentation","url":"/task/video-object-segmentation","datasets_with_task":"/datasets/task/video-object-segmentation"},{"name":"Video Understanding","url":"/task/video-understanding","datasets_with_task":"/datasets/task/video-understanding"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["M$^3$-VOS"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/video-object-segmentation-on-m-3-vos","task":"Video Object Segmentation","dataset_variant":"M$^3$-VOS","rows":4,"metrics":["Average IOU"],"first_row_in_archive_order":{"model":"ReVOS","paper":"/paper/m-3-vos-multi-phase-multi-transition-and-1","metrics":{"Average IOU":"75.6"},"code_links":[{"title":"zixuan-chen/M3VOS_Experiment","url":"https://github.com/zixuan-chen/M3VOS_Experiment"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/m-3-vos-multi-phase-multi-transition-and-1","title":"M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation","date":"2025-06-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/2408-00714","title":"SAM 2: Segment Anything in Images and Videos","date":"2024-08-01","rows_on_this_dataset":1,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":49,"samples_ran":28,"samples_unverified":21,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/putting-the-object-back-into-video-object","title":"Putting the Object Back into Video Object Segmentation","date":"2023-10-19","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":3,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/xmem-long-term-video-object-segmentation-with","title":"XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model","date":"2022-07-14","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":1,"samples_unverified":2,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":57,"samples_ran":32,"samples_unverified":25,"pointer_only_for_licence":2,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}