| From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering added by Syntology |
2026-03 (from id) |
VILA-Lab/PIXAR/model/PIXAR.py b9e60254a36a3fb8 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| UGround: Towards Unified Visual Grounding with Unrolled Transformers added by Syntology |
2025-10 (from id) |
rui-qian/UGround/model/UGround.py 3face96ac23ae5bc |
unverified |
Apache-2.0 (permissive) |
| UGround: Towards Unified Visual Grounding with Unrolled Transformers added by Syntology |
2025-10 (from id) |
rui-qian/UGround/model/GSVA.py 13c0bb91f8009d9c |
unverified |
Apache-2.0 (permissive) |
| Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation added by Syntology |
2025-08 (from id) |
JiahuaDong/HVPL/hvpl/modeling/vita_criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation |
17 Jul 2025 |
HuangShiqi128/SCORE/score/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion |
11 Jul 2025 |
Enyu-Liu/DISC/maskdino/models/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning |
27 Jun 2025 |
geshang777/FOCUS/focus/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| Towards In-the-wild 3D Plane Reconstruction from a Single Image |
3 Jun 2025 |
jcliu0428/ZeroPlane/ZeroPlane/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing |
23 Jan 2025 |
mbzuai-oryx/geopixel/model/geopixel.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model |
5 Dec 2024 |
identical code first harvested elsewhere a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation |
4 Dec 2024 |
ruohaoguo/avis/avism/modeling/avism_criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| A Simple Image Segmentation Framework via In-Context Examples |
7 Oct 2024 |
aim-uofa/SINE/sine/model/criterion.py 7130af6c2979dba8 |
ran
fingerprinted |
licence not identified · pointer only |
| One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos |
29 Sep 2024 |
showlab/videolisa/model/VideoLISA.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects |
23 Jul 2024 |
ProvenceStar/PartGLEE/projects/PartGLEE/partglee/models/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| VISA: Reasoning Video Object Segmentation via Large Language Models |
16 Jul 2024 |
cilinyan/VISA/model/VISA.py bf7e8883d6ce76b4 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Part2Object: Hierarchical Unsupervised 3D Instance Segmentation |
14 Jul 2024 |
chengshiest/part2object/models/criterion.py c0938fc9bc7ebddd |
ran
fingerprinted |
MIT (permissive) |
| EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model |
28 Jun 2024 |
identical code first harvested elsewhere a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Open-Vocabulary Semantic Segmentation with Image Embedding Balancing |
14 Jun 2024 |
slonetime/EBSeg/ebseg/model/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Instruction-Guided Visual Masking |
30 May 2024 |
2toinf/ivm/model/IVM.py 1033e9116b1e2573 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Grounded 3D-LLM with Referent Tokens |
16 May 2024 |
OpenRobotLab/Grounded_3D-LLM/models/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference |
9 May 2024 |
lzhxmu/vtw/LISA-VTW/LISA.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation |
23 Apr 2024 |
abhishekaich27/proscale-pytorch/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| LaSagnA: Language-based Segmentation Assistant for Complex Queries |
12 Apr 2024 |
congvvc/lasagna/model/LaSagnA.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| CoReS: Orchestrating the Dance of Reasoning and Segmentation |
8 Apr 2024 |
identical code first harvested elsewhere a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation |
4 Apr 2024 |
heshuting555/DsHmp/dshmp/modeling/vita_criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception |
5 Mar 2024 |
jwh97nn/AnyRef/model/anyref.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| PEM: Prototype-based Efficient MaskFormer for Image Segmentation |
29 Feb 2024 |
niccolocavagnero/pem/pem/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| UniVS: Unified and Universal Video Segmentation with Prompts as Queries |
28 Feb 2024 |
minghanli/univs/univs/modeling/video_criterion.py 450b62615caa903a |
ran
fingerprinted |
no licence file found · pointer only |
| UniVS: Unified and Universal Video Segmentation with Prompts as Queries |
28 Feb 2024 |
minghanli/univs/univs/modeling/video_criterion_prompt.py 46b849406325dd41 |
ran
fingerprinted |
no licence file found · pointer only |
| Symbol as Points: Panoptic Symbol Spotting via Point-based Representation |
19 Jan 2024 |
nicehuster/sympoint/svgnet/model/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Supervised Fine-tuning in turn Improves Visual Foundation Models |
18 Jan 2024 |
tencentarc/visft/mmf/models/visft/segment_criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| ODIN: A Single Model for 2D and 3D Segmentation |
4 Jan 2024 |
ayushjain1144/odin/odin/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model |
28 Dec 2023 |
dvlab-research/lisa/model/LISA.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces |
25 Dec 2023 |
foundationvision/uniref/projects/UniRef/uniref/models/uniref_sam.py e28a878991e1e696 |
ran
fingerprinted |
MIT (permissive) |
| GSVA: Generalized Segmentation via Multimodal Large Language Models |
15 Dec 2023 |
leaplabthu/gsva/model/losses.py 13c0bb91f8009d9c |
unverified |
Apache-2.0 (permissive) |
| Collaborating Foundation Models for Domain Generalized Semantic Segmentation |
15 Dec 2023 |
yasserben/clouds/clouds/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| General Object Foundation Model for Images and Videos at Scale |
14 Dec 2023 |
FoundationVision/GLEE/projects/GLEE/glee/models/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation |
11 Dec 2023 |
rkzheng99/TMT-VIS/tmt/modeling/tmt_criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Open-Vocabulary Segmentation with Semantic-Assisted Calibration |
7 Dec 2023 |
workforai/scan/scan/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| PixelLM: Pixel Reasoning with Large Multimodal Model |
4 Dec 2023 |
maverickren/pixellm/model/PixelLM.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Mask Propagation for Efficient Video Semantic Segmentation |
29 Oct 2023 |
ziplab/MPVSS/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
no licence file found · pointer only |
| Learning Mask-aware CLIP Representations for Zero-Shot Segmentation |
30 Sep 2023 |
jiaosiyu1999/MAFT-Plus/maft/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| 3D Indoor Instance Segmentation in an Open-World |
25 Sep 2023 |
aminebdj/3D-OWIS/models/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Detect Everything with Few Examples |
22 Sep 2023 |
mlzxy/devit/detectron2/modeling/meta_arch/devit_update.py 847c0c0c48d84c9e |
ran
fingerprinted |
MIT (permissive) |
| MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions |
16 Aug 2023 |
henghuiding/MeViS/lmpm/modeling/vita_criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning |
12 Aug 2023 |
aim-uofa/segprompt/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
BSD-2-Clause (permissive) |
| Symphonize 3D Semantic Scene Completion with Contextual Instance Queries |
27 Jun 2023 |
hustvl/symphonies/maskdino/models/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation |
19 Jun 2023 |
heshuting555/PADing/PADing/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Segment Any Point Cloud Sequences by Distilling Vision Foundation Models |
15 Jun 2023 |
IDEA-Research/OpenSeeD/openseed/modules/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| Compositor: Bottom-up Clustering and Compositing for Robust Part and Object Segmentation |
12 Jun 2023 |
tacju/compositor/Compositor_Mask2Former/compositor/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| DFormer: Diffusion-guided Transformer for Universal Image Segmentation |
6 Jun 2023 |
cp3wan/dformer/dformer/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| AdaptiveClick: Clicks-aware Transformer with Adaptive Focal Loss for Interactive Image Segmentation |
7 May 2023 |
lab206/adaptiveclick/isegm/model/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Complementary Random Masking for RGB-Thermal Semantic Segmentation |
30 Mar 2023 |
UkcheolShin/CRM_RGBTSeg/models/mask2former/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| You Only Segment Once: Towards Real-Time Panoptic Segmentation |
26 Mar 2023 |
Darth-Kronos/YOSO_TensorRT/projects/YOSO/yoso/loss.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| BoxSnake: Polygonal Instance Segmentation with Box Supervision |
21 Mar 2023 |
Yangr116/BoxSnake/modeling/box_supervisor/box_supervisor.py 18da425e8ba54fbd |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation |
15 Mar 2023 |
junjiehe96/fastinst/fastinst/modeling/criterion.py a2b19dbf3e1969b9 |
unverified |
MIT (permissive) |
| Side Adapter Network for Open-Vocabulary Semantic Segmentation |
23 Feb 2023 |
blumenstiel/SAN-MESS/san/model/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| OneFormer: One Transformer to Rule Universal Image Segmentation |
10 Nov 2022 |
SHI-Labs/OneFormer/oneformer/modeling/criterion.py dc4a40c394479bd8 |
unverified |
MIT (permissive) |
| Mask3D: Mask Transformer for 3D Semantic Instance Segmentation |
6 Oct 2022 |
jonasschult/mask3d/models/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| VITA: Video Instance Segmentation via Object Token Association |
9 Jun 2022 |
sukjunhwang/vita/vita/modeling/vita_criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| RankSeg: Adaptive Pixel Classification with Image Category Ranking for Segmentation |
8 Mar 2022 |
facebookresearch/Mask2Former/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection |
7 Mar 2022 |
IDEACVR/MaskDINO/maskdino/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| Mask2Former for Video Instance Segmentation |
20 Dec 2021 |
nihalsid/mask2former/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| Masked-attention Mask Transformer for Universal Image Segmentation |
2 Dec 2021 |
DdeGeus/Mask2Former-IBS/mask2former/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| arXiv:aaai_27844 |
|
retsuh-bqw/SRFormer-Text-Det/adet/modeling/srformer/matcher.py 641305cea54a6936 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Zhou_LIRA_Reasoning_Reconstruction_via_Multimodal_Large_Language_Models_ICCV_2025_paper |
|
zhen6618/LIRA/LIRA/models/LISA.py a9292f5d89194794 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:Zhang_Uni-3D_A_Universal_Model_for_Panoptic_3D_Scene_Reconstruction_ICCV_2023_paper |
|
mlpc-ucsd/Uni-3D/uni_3d/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Zhang_FreePoint_Unsupervised_Point_Cloud_Instance_Segmentation_CVPR_2024_paper |
|
zzk273/FreePoint/models/criterion_freepoint.py d0c61e8dba511aa3 |
unverified |
MIT (permissive) |
| arXiv:Xia_GSVA_Generalized_Segmentation_via_Multimodal_Large_Language_Models_CVPR_2024_paper |
|
LeapLabTHU/GSVA/model/losses.py 13c0bb91f8009d9c |
unverified |
Apache-2.0 (permissive) |
| arXiv:Qian_Reasoning_to_Attend_Try_to_Understand_How_SEG_Token_Works_CVPR_2025_paper |
|
rui-qian/READ/model/READ.py 3face96ac23ae5bc |
unverified |
MIT (permissive) |
| arXiv:Li_Mask_DINO_Towards_a_Unified_Transformer-Based_Framework_for_Object_Detection_CVPR_2023_paper |
|
IDEA-Research/MaskDINO/maskdino/modeling/criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Heo_A_Generalized_Framework_for_Video_Instance_Segmentation_CVPR_2023_paper |
|
miranheo/GenVIS/genvis/modeling/genvis_criterion.py d0c61e8dba511aa3 |
unverified |
Apache-2.0 (permissive) |