| CHARTOGRAPHER: Counterfactual Chart Generation for Evaluating Vision-Language Models added by Syntology |
2026-05 (from id) |
vis-nlp/ChartQA/Models/VL-T5/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
GPL-3.0 (copyleft) · pointer only |
| Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation |
6 Jun 2025 |
liyih/CSCL/code/MultiModal-DeepFake-main/models/METER/dist_utils.py 53838c15d969eb79 |
ran
|
MIT (permissive) |
| ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement |
2 Apr 2025 |
illume-unified-mllm/ILLUME_plus/ILLUME/illume/dist_utils.py 2a3280757ce5a198 |
ran · violated contract
|
Apache-2.0 (permissive) |
| CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification |
19 Mar 2025 |
YuWLong666/CoE/util/misc.py 05992505df9ef0cd |
ran
|
BSD-3-Clause (permissive) |
| Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning |
17 Dec 2024 |
ShipingGe/ILCACM/src/util/dist.py a38ed1aa8d59134e |
unverified |
MIT (permissive) |
| D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement |
17 Oct 2024 |
Peterande/D-FINE/src/misc/logger.py e6c2f1b16c9d4813 |
unverified |
Apache-2.0 (permissive) |
| A Simple Image Segmentation Framework via In-Context Examples |
7 Oct 2024 |
aim-uofa/SINE/sine/utils/comm.py 53838c15d969eb79 |
ran
|
licence not identified · pointer only |
| X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation |
28 Sep 2024 |
pinxueguo/x-prompt/lib/utils/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| MeshAnything V2: Artist-Created Mesh Generation With Adjacent Mesh Tokenization |
5 Aug 2024 |
buaacyw/meshanythingv2/meshanything_train/dist.py b420987753130180 |
ran
|
licence not identified · pointer only |
| SAM 2: Segment Anything in Images and Videos |
1 Aug 2024 |
louisfinner/him2sam/lib/utils/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| LLaRA: Supercharging Robot Learning Data for Vision-Language Policy |
28 Jun 2024 |
LostXine/LLaRA/maskrcnn/utils.py d4ff7353113eadb2 |
ran
|
Apache-2.0 (permissive) |
| SEMv3: A Fast and Robust Approach to Table Separation Line Detection |
20 May 2024 |
Chunchunwumu/SEMv3/cal_teds/utils.py 2a3280757ce5a198 |
ran · violated contract
|
Apache-2.0 (permissive) |
| Selective Focus: Investigating Semantics Sensitivity in Post-training Quantization for Lane Detection |
10 May 2024 |
PannenetsF/SelectiveFocus/pad/utils/runners/lane_det_quant_trainer.py 3a37f431350cb887 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Sketch-guided Image Inpainting with Partial Discrete Diffusion Process |
18 Apr 2024 |
vl2g/sketch-inpainting/image_synthesis/distributed/distributed.py 7dbdda0920e3d074 |
ran
|
MIT (permissive) |
| ClusterTabNet: Supervised clustering method for table detection and table structure recognition |
12 Feb 2024 |
sap-samples/clustertabnet/train/util.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models |
24 Jan 2024 |
hkbunlp/explainhm-www2024/src/modules/dist_utils.py 53838c15d969eb79 |
ran
|
Apache-2.0 (permissive) |
| PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation |
23 Jan 2024 |
xzz2/pa-sam/utils/misc.py 4183d276c611a490 |
ran
|
no licence file found · pointer only |
| GPAvatar: Generalizable and Precise Head Avatar from Image(s) |
18 Jan 2024 |
xg-chu/gpavatar/core/utils/distributed.py 53838c15d969eb79 |
ran
|
MIT (permissive) |
| UMIE: Unified Multimodal Information Extraction with Instruction Tuning |
5 Jan 2024 |
ZUCC-AI/UMIE/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
no licence file found · pointer only |
| Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection |
27 Dec 2023 |
Michel-liu/FatFormer/utils/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| Exploring news intent and its application: A theory-driven approach |
27 Dec 2023 |
ictmcg/newsint/DMint/utils/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection |
28 Nov 2023 |
easonxiao-888/uvcom/uvcom/misc_ddp.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| Unleashing the Power of Prompt-driven Nucleus Instance Segmentation |
27 Nov 2023 |
windygoo/promptnucseg/prompter/utils.py 0f432e22853a1915 |
ran
|
no licence file found · pointer only |
| SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting |
16 Nov 2023 |
hcplab-sysu/sqlnet/util/utils.py 8ae0b62e37947d8e |
unverified |
no licence file found · pointer only |
| CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation |
18 Sep 2023 |
aspirinone/catr.github.io/CATR/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| LAN-HDR: Luminance-based Alignment Network for High Dynamic Range Video Reconstruction |
22 Aug 2023 |
haesoochung/lan-hdr/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models |
18 Aug 2023 |
mlvlab/OVQA/util/dist.py a38ed1aa8d59134e |
unverified |
no licence file found · pointer only |
| RestoreFormer++: Towards Real-World Blind Face Restoration from Undegraded Key-Value Pairs |
14 Aug 2023 |
wzhouxiff/restoreformerplusplus/RestoreFormer/distributed/distributed.py 7dbdda0920e3d074 |
ran
|
Apache-2.0 (permissive) |
| Random Boxes Are Open-world Object Detectors |
17 Jul 2023 |
scuwyh2000/RandBox/randbox/util/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering |
29 Jun 2023 |
Abhiram4572/mi_bart/mi_bart/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| OpenP5: An Open-Source Platform for Developing, Training, and Evaluating LLM-based Recommender Systems |
19 Jun 2023 |
jeykigung/P5/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| MixFormerV2: Efficient Fully Transformer Tracking |
25 May 2023 |
mcg-nju/mixformerv2/lib/utils/misc.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| VIP5: Towards Multimodal Foundation Models for Recommendation |
23 May 2023 |
jeykigung/vip5/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| What Makes for Good Visual Tokenizers for Large Language Models? |
20 May 2023 |
tencentarc/gvt/gvt/gvt/modules/dist_utils.py 53838c15d969eb79 |
ran
|
Apache-2.0 (permissive) |
| Hierarchical Video-Moment Retrieval and Step-Captioning |
29 Mar 2023 |
j-min/HiREST/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| The Devil is in the Points: Weakly Semi-Supervised Instance Segmentation via Point-Guided Mask Representation |
27 Mar 2023 |
clovaai/PointWSSIS/MaskRefineNet/utils/distributed.py 426f9359bfa8b5d4 |
unverified |
Apache-2.0 (permissive) |
| You Only Segment Once: Towards Real-Time Panoptic Segmentation |
26 Mar 2023 |
Darth-Kronos/YOSO_TensorRT/projects/YOSO/yoso/utils.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| Weakly Supervised Knowledge Transfer with Probabilistic Logical Reasoning for Object Detection |
9 Mar 2023 |
molden/ProbKT/robust_detection/dpl_utils.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| SPTS v2: Single-Point Scene Text Spotting |
4 Jan 2023 |
bytedance/sptsv2/util/misc_sptsv2.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| DiffusionInst: Diffusion Model for Instance Segmentation |
6 Dec 2022 |
alipay/diffusion-model-for-instance-segmentation/diffusioninst/util/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| Cross-Modal Adapter for Text-Video Retrieval |
17 Nov 2022 |
leaplabthu/cross-modal-adapter/comm.py 2a3280757ce5a198 |
ran · violated contract
|
Apache-2.0 (permissive) |
| Multimedia Generative Script Learning for Task Planning |
25 Aug 2022 |
EagleW/Multimedia-Generative-Script-Learning-for-Task-Planning/utils_.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement |
18 Aug 2022 |
zhiqic/gsrformer/util/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Occupancy Planes for Single-view RGB-D Human Reconstruction |
4 Aug 2022 |
xiaoming-zhao/oplanes/oplanes/utils/comm.py 53838c15d969eb79 |
ran
|
Apache-2.0 (permissive) |
| ReAct: Temporal Action Detection with Relational Queries |
14 Jul 2022 |
sssste/React/React/utill/misc.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| Improved Vector Quantized Diffusion Models |
31 May 2022 |
microsoft/vq-diffusion/image_synthesis/distributed/distributed.py 7dbdda0920e3d074 |
ran
|
MIT (permissive) |
| On the Paradox of Learning to Reason from Data |
23 May 2022 |
joshuacnf/paradox-learning2reason/dist.py a38ed1aa8d59134e |
unverified |
MIT (permissive) |
| Semi-Supervised Training to Improve Player and Ball Detection in Soccer |
14 Apr 2022 |
rvandeghen/sst/src/utils.py 05992505df9ef0cd |
ran
|
BSD-3-Clause (permissive) |
| Disentangled Representation Learning for Text-Video Retrieval |
14 Mar 2022 |
foolwood/DRL/tvr/utils/comm.py 2a3280757ce5a198 |
ran · violated contract
|
Apache-2.0 (permissive) |
| Rethinking Efficient Lane Detection via Curve Modeling |
4 Mar 2022 |
voldemortX/pytorch-auto-drive/utils/ddp_utils.py d2092a95e79d2d3b |
unverified |
BSD-3-Clause (permissive) |
| Rethinking the Two-Stage Framework for Grounded Situation Recognition |
10 Dec 2021 |
kellyiss/situformer/util/misc.py 05992505df9ef0cd |
ran
|
no licence file found · pointer only |
| End-to-End Referring Video Object Segmentation with Multimodal Transformers |
29 Nov 2021 |
mttr2021/MTTR/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| Grounded Situation Recognition with Transformers |
19 Nov 2021 |
jhcho99/gsrtr/util/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Elaborative Rehearsal for Zero-shot Action Recognition |
5 Aug 2021 |
DeLightCMU/ElaborativeRehearsal/framework/dist_helper.py 5b2ec24a42ad7b48 |
unverified |
MIT (permissive) |
| TransVG: End-to-End Visual Grounding with Transformers |
17 Apr 2021 |
nku-shengzheliu/Pytorch-TransVG/utils/misc.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| TubeR: Tubelet Transformer for Video Action Detection |
2 Apr 2021 |
amazon-science/tubelet-transformer/models/transformer/util/misc.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| CvT: Introducing Convolutions to Vision Transformers |
29 Mar 2021 |
microsoft/CvT/lib/utils/comm.py 19a8d1208f8a920f |
unverified |
MIT (permissive) |
| Attention-based Joint Detection of Object and Semantic Part |
5 Jul 2020 |
kevalmorabia97/Object-and-Semantic-Part-Detection-pyTorch/references/detection/utils.py 05992505df9ef0cd |
ran
|
Apache-2.0 (permissive) |
| Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild |
4 Apr 2020 |
EAST-J/Youtubehand/utils/comm.py 2a3280757ce5a198 |
ran · violated contract
|
MIT (permissive) |
| Learning Delicate Local Representations for Multi-Person Pose Estimation |
9 Mar 2020 |
caiyuanhao1998/RSN/lib/utils/comm.py 2a3280757ce5a198 |
ran · violated contract
|
MIT (permissive) |
| Zero-Shot Grounding of Objects from Natural Language Queries |
20 Aug 2019 |
TheShadow29/zsgnet-pytorch/code/utils.py b7654b7006fe061d |
unverified |
MIT (permissive) |
| IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments |
26 Nov 2018 |
prajjwal1/autonomous-object-detection/utils.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| ATOM: Accurate Tracking by Overlap Maximization |
19 Nov 2018 |
xuefeng-zhu5/cdaat/lib/utils/misc.py 5c8f207cb5f19197 |
unverified |
MIT (permissive) |
| arXiv:aaai_25418 |
|
Shinetism/VStates/utils/trn_utils.py ef34b694a1a97e7e |
unverified |
MIT (permissive) |
| arXiv:Zhang_VQACL_A_Novel_Visual_Question_Answering_Continual_Learning_Setting_CVPR_2023_paper |
|
zhangxi1997/VQACL/VL-T5/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| arXiv:Yuan_CAT_A_Unified_Click-and-Track_Framework_for_Realistic_Tracking_ICCV_2025_paper |
|
ysyuann/CAT/lib/utils/misc.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| arXiv:Yoshiyasu_Deformable_Mesh_Transformer_for_3D_Human_Mesh_Recovery_CVPR_2023_paper |
|
yusukey03012/DeFormer/src/utils/comm.py 2a3280757ce5a198 |
ran · violated contract
|
MIT (permissive) |
| arXiv:Xiao_Bridging_the_Gap_A_Unified_Video_Comprehension_Framework_for_Moment_CVPR_2024_paper |
|
EasonXiao-888/UVCOM/uvcom/misc_ddp.py 05992505df9ef0cd |
ran
|
MIT (permissive) |
| arXiv:Wang_Object_Detection_using_Event_Camera_A_MoE_Heat_Conduction_based_CVPR_2025_paper |
|
Event-AHU/OpenEvDET/CvHeat-DET/src/misc/logger.py e6c2f1b16c9d4813 |
unverified |
MIT (permissive) |
| arXiv:Gu_Vector_Quantized_Diffusion_Model_for_Text-to-Image_Synthesis_CVPR_2022_paper |
|
cientgu/VQ-Diffusion/image_synthesis/distributed/distributed.py 7dbdda0920e3d074 |
ran
|
MIT (permissive) |
| arXiv:2023.findings-emnlp.644 |
|
jeykigung/VIP5/src/dist_utils.py 44f4152cc7eb1ddd |
ran
|
MIT (permissive) |
| arXiv:136730727 |
|
louisYen/S3R/anomaly/apis/comm.py 2a3280757ce5a198 |
ran · violated contract
|
MIT (permissive) |
| arXiv:136640328 |
|
nutuniv/SSRL/misc.py 05992505df9ef0cd |
ran
|
MIT (permissive) |