| Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language added by Syntology |
2026-09 (from id) |
KIT-MRT/AD-Diff/serve/utils_general.py dd294474d2020a4c |
unverified |
Apache-2.0 (permissive) |
| WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks added by Syntology |
2026-04 (from id) |
wi-pi/webspeval_code/src/utils.py e63c4f0834fe9c79 |
unverified |
Apache-2.0 (permissive) |
| RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies added by Syntology |
2026-03 (from id) |
RoboMME/MemoryVLA/deploy.py ed5a57b34c0d97f3 |
unverified |
no licence file found · pointer only |
| Stylizing ViT: Anatomy-Preserving Instance Style Transfer for Domain Generalization added by Syntology |
2026-01 (from id) |
sdoerrich97/stylizing-vit/stylizing_vit/util.py 4b10e9d97b27d49a |
unverified |
Apache-2.0 (permissive) |
| ImLoc: Revisiting Visual Localization with Image-based Representation added by Syntology |
2026-01 (from id) |
cvg/Hierarchical-Localization/hloc/extract_features.py 47ffaeadde70c45c |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.09168 |
2025-07 (from id) |
Alex-Zhu1/SSD/ui_utils.py cbb2f431f7f749fc |
unverified |
licence not identified · pointer only |
| HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer |
28 May 2025 |
hidream-ai/hidream-e1/inference_e1_1.py 52252ec23d3c110c |
ran · our draft was wrong
|
MIT (permissive) |
| Re-Evaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning Model |
2 Mar 2025 |
rundonghe/RESSL/evaluation_unseen_near_change.py cf297d9b65ad7f32 |
unverified |
MIT (permissive) |
| Magma: A Foundation Model for Multimodal AI Agents |
18 Feb 2025 |
microsoft/Magma/agents/libero/libero_env_utils.py 090fa82a9a0f5b4b |
unverified |
MIT (permissive) |
| Aria-UI: Visual Grounding for GUI Instructions |
20 Dec 2024 |
ariaui/aria-ui/utils.py 6d4730f92726b8ee |
ran · our draft was wrong
|
no licence file found · pointer only |
| Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning |
14 Dec 2024 |
heimingx/tag/eval_mm/screenspot/eval_TAG.py 70dc9b613a4d20df |
ran · fixture could not drive it
|
no licence file found · pointer only |
| VisionZip: Longer is Better but Not Necessary in Vision Language Models |
5 Dec 2024 |
dvlab-research/visionzip/gradio_demo.py f170144b6e9abcec |
unverified |
Apache-2.0 (permissive) |
| AnyText2: Visual Text Generation and Editing With Customizable Attributes |
22 Nov 2024 |
tyxsspa/anytext2/util.py a91c692ff279a01a |
unverified |
Apache-2.0 (permissive) |
| FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on |
15 Nov 2024 |
BoyuanJiang/FitDiT/gradio_sd3.py 9c6f71ed6f24968f |
ran · our draft was wrong
|
licence not identified · pointer only |
| Structure Consistent Gaussian Splatting with Matching Prior for Few-shot Novel View Synthesis |
6 Nov 2024 |
prstrive/scgaussian/data_preprocess/get_match_info.py 55ec6eda769b51be |
unverified |
licence not identified · pointer only |
| Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image Synthesis |
29 Oct 2024 |
DeepakSridhar/fgdm/controlnet/annotator/util.py 50fd68f6989503c6 |
unverified |
licence not identified · pointer only |
| OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization |
25 Oct 2024 |
minorjerry/openwebvoyager/WebVoyager/utils.py e63c4f0834fe9c79 |
unverified |
Apache-2.0 (permissive) |
| MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models |
23 Oct 2024 |
liuziyu77/mia-dpo/LLaVA-Hound-DPO/chatuniv/chatuniv_utils.py c3f1f1085b84693b |
unverified |
Apache-2.0 (permissive) |
| MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems |
18 Oct 2024 |
zivenzhu/multi-chart-qa/code/evaluate_gemini.py 3f22b07dd2f94a8c |
ran
|
licence not identified · pointer only |
| VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding |
17 Oct 2024 |
openrobotlab/vlm-grounder/vlm_grounder/utils/grid_image_generator.py ea5a2af68864cf1d |
ran
|
no licence file found · pointer only |
| Law of the Weakest Link: Cross Capabilities of Large Language Models |
30 Sep 2024 |
facebookresearch/llm-cross-capabilities/evaluation/utils.py b4878d3dddc25ebb |
ran
|
licence not identified · pointer only |
| Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding |
5 Sep 2024 |
yunzeman/lexicon3d/lexicon3d/models.py 2a9ddb0649c62fe4 |
unverified |
MIT (permissive) |
| Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation |
19 Aug 2024 |
yuyangye/mllm-msr/MLLM-MSR/train/microlens/train_llava_sft.py d845ce79b34089e6 |
ran
|
no licence file found · pointer only |
| Towards Reliable Advertising Image Generation Using Human Feedback |
1 Aug 2024 |
ZhenbangDu/Reliable_AD/utils/utils.py e14248108fd2691b |
ran
fingerprinted |
no licence file found · pointer only |
| DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation |
16 Jul 2024 |
kaist-cvml-lab/DreamCatalyst/nerfstudio/dc/dc.py 625f7096a79b118d |
ran
|
no licence file found · pointer only |
| DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation |
16 Jul 2024 |
kaist-cvml-lab/DreamCatalyst/threestudio/ui_utils.py cbb2f431f7f749fc |
unverified |
no licence file found · pointer only |
| StyleShot: A Snapshot on Any Style |
1 Jul 2024 |
open-mmlab/StyleShot/annotator/util.py 50fd68f6989503c6 |
unverified |
MIT (permissive) |
| AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation |
27 Jun 2024 |
open-mmlab/anycontrol/annotator/util.py 50fd68f6989503c6 |
unverified |
MIT (permissive) |
| Learning to Continually Learn with the Bayesian Principle |
29 May 2024 |
soochan-lee/SB-MCL/dataset.py 20f1b95de7047461 |
ran
|
no licence file found · pointer only |
| TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing |
17 Apr 2024 |
sherryxtchen/tino-edit/utils.py b360684f1ea3a80d |
ran
|
Apache-2.0 (permissive) |
| MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control |
18 Mar 2024 |
Zhoues/MineDreamer/minedreamer/VPT/agent.py 7191618c00cc9152 |
ran
|
Apache-2.0 (permissive) |
| CRS-Diff: Controllable Remote Sensing Image Generation with Diffusion Model |
18 Mar 2024 |
Sonettoo/CRS-Diff/annotator/util.py 50fd68f6989503c6 |
unverified |
no licence file found · pointer only |
| NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear Interpolation |
13 Mar 2024 |
tmlr-group/NoiseDiffusion/controlnet/annotator/util.py 50fd68f6989503c6 |
unverified |
no licence file found · pointer only |
| DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations |
11 Mar 2024 |
bytedance/deadiff/ldm/controlnet/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |
| What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? |
10 Mar 2024 |
aim-uofa/genpercept/GenPercept_v1/hubconf.py 50fd68f6989503c6 |
unverified |
BSD-2-Clause (permissive) |
| GIM: Learning Generalizable Image Matcher From Internet Videos |
16 Feb 2024 |
xuelunshen/gim/hloc/extract_features.py 47ffaeadde70c45c |
unverified |
MIT (permissive) |
| I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models |
27 Dec 2023 |
KwaiVGI/I2V-Adapter/src/utils/util.py 2e3e408dc6b5d79d |
ran
fingerprinted |
no licence file found · pointer only |
| Towards Real-World Blind Face Restoration with Generative Diffusion Prior |
25 Dec 2023 |
chenxx89/bfrffusion/annotator/util.py 50fd68f6989503c6 |
unverified |
MIT (permissive) |
| ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems |
11 Dec 2023 |
vislearn/ControlNet-XS/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |
| BEDD: The MineRL BASALT Evaluation and Demonstrations Dataset for Training and Benchmarking Agents that Solve Fuzzy Tasks |
5 Dec 2023 |
minerllabs/basalt-benchmark/basalt/vpt_lib/agent.py 69a190b789185068 |
ran
|
MIT (permissive) |
| GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting |
24 Nov 2023 |
buaacyw/gaussianeditor/ui_utils.py cbb2f431f7f749fc |
unverified |
no licence file found · pointer only |
| RoboDepth: Robust Out-of-Distribution Depth Estimation under Corruptions |
23 Oct 2023 |
isl-org/DPT/util/io.py 80ea105639cd9efd |
ran
|
MIT (permissive) |
| CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulation |
19 Oct 2023 |
sled-group/cyclenet/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |
| Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis |
7 Sep 2023 |
zhujiapeng/aurora/utils/image_utils.py 93a805b315b09a0b |
unverified |
licence not identified · pointer only |
| StableVideo: Text-driven Consistency-aware Diffusion Video Editing |
18 Aug 2023 |
rese1f/stablevideo/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |
| LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding |
29 Jun 2023 |
SALT-NLP/LLaVAR/LLaVA/llava/eval/model_vqa.py 0c495b53fac1068b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond |
26 May 2023 |
thu-ml/controlvideo/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |
| Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models |
25 May 2023 |
shihaozhaozsh/uni-controlnet/annotator/util.py 50fd68f6989503c6 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| Towards Solving Fuzzy Tasks with Human Feedback: A Retrospective of the MineRL BASALT 2022 Competition |
23 Mar 2023 |
shuishida/minerl_2022/openai_vpt/agent.py 69a190b789185068 |
ran
|
MIT (permissive) |
| Visual Language Maps for Robot Navigation |
11 Oct 2022 |
vlmaps/vlmaps/vlmaps/lseg/additional_utils/encoding_models.py 2a9ddb0649c62fe4 |
unverified |
MIT (permissive) |
| Attention Consistency on Visual Corruptions for Single-Source Domain Generalization |
27 Apr 2022 |
explainableml/acvc/tools.py af479d88352e71eb |
unverified |
MIT (permissive) |
| Document Dewarping with Control Points |
20 Mar 2022 |
gwxie/document-dewarping-with-control-points/Source/dataloader.py 477294970cbee196 |
unverified |
MIT (permissive) |
| Low-Rank Subspaces in GANs |
8 Jun 2021 |
zhujiapeng/LowRankGAN/compute_jacobian.py 29bfeffdb65e32bc |
unverified |
MIT (permissive) |
| MobileDets: Searching for Object Detection Architectures for Mobile Accelerators |
30 Apr 2020 |
inacmor/mobiledets-yolov4-pytorch/dataset/datasets.py cfd33b179e7cf880 |
unverified |
MIT (permissive) |
| YOLOv4: Optimal Speed and Accuracy of Object Detection |
23 Apr 2020 |
bubbliiiing/yolov4-pytorch/utils/utils.py 43bbe5f830d240ce |
unverified |
MIT (permissive) |
| YOLOv4: Optimal Speed and Accuracy of Object Detection |
23 Apr 2020 |
hhk7734/tensorflow-yolov4/py_src/yolov4/common/media.py 51ff7bacc4b85623 |
unverified |
MIT (permissive) |
| FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking |
4 Apr 2020 |
harsh2912/people-tracking/src/lib/tracking_utils/visualization.py 15af20c349fec967 |
unverified |
MIT (permissive) |
| Real-time Scene Text Detection with Differentiable Binarization |
20 Nov 2019 |
Mushroomcat9998/DBNet/data_loader/modules/augment.py 74c3e5afa7495d81 |
unverified |
Apache-2.0 (permissive) |
| Spatiotemporal Tile-based Attention-guided LSTMs for Traffic Video Prediction |
24 Oct 2019 |
tumeteor/neurips2019challenge/src/data_utils/loader.py 46c330ffbb98dc2b |
unverified |
Apache-2.0 (permissive) |
| Boundary-Aware Feature Propagation for Scene Segmentation |
31 Aug 2019 |
henghuiding/BFP/utils/models/base.py 953a5c096d03f6f4 |
unverified |
MIT (permissive) |
| Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer |
2 Jul 2019 |
anlok/depthmap-loktev/utils.py 80ea105639cd9efd |
ran
|
MIT (permissive) |
| Learnable Triangulation of Human Pose |
14 May 2019 |
karfly/learnable-triangulation-pytorch/mvn/utils/img.py 3658ef58769f24d4 |
unverified |
MIT (permissive) |
| RetinaFace: Single-stage Dense Face Localisation in the Wild |
2 May 2019 |
serengil/deepface/deepface/modules/preprocessing.py f3ebcbef63b38e8b |
unverified |
MIT (permissive) |
| Gradient-free activation maximization for identifying effective stimuli |
1 May 2019 |
willwx/XDream/xdream/net_utils/transformer.py 74848d8128c1942e |
unverified |
MIT (permissive) |
| Gradient-free activation maximization for identifying effective stimuli |
1 May 2019 |
willwx/XDream/xdream/utils.py 544047023db0e72f |
unverified |
MIT (permissive) |
| From Recognition to Cognition: Visual Commonsense Reasoning |
27 Nov 2018 |
karanchahal/play-with-vcr/dataloaders/box_utils.py 086cec07d5d5e16d |
unverified |
MIT (permissive) |
| Self-Attention Generative Adversarial Networks |
21 May 2018 |
vijishmadhavan/SkinDeep/tattoorem.py 9519c43c0017c576 |
unverified |
Apache-2.0 (permissive) |
| MGGAN: Solving Mode Collapse using Manifold Guided Training |
12 Apr 2018 |
QuickSolverKyle/Tensorflow-MyGANs/utils.py d0a4f3b63424b46e |
unverified |
MIT (permissive) |
| COCO-Stuff: Thing and Stuff Classes in Context |
12 Dec 2016 |
divyanshpuri02/COCO_2018-Stuff-Segmentation-Challenge/COCO_2018-Stuff-Segmentation-Challenge/keras_segmentation/models/model_utils.py e644ee2441042518 |
unverified |
MIT (permissive) |
| A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering |
23 Nov 2016 |
totalgood/viddesc/src/imgdesc/resize.py ddccddb3e642d270 |
unverified |
MIT (permissive) |
| Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization |
7 Oct 2016 |
TooTouch/WhiteBox-Part1/code/saliency/attribution_methods.py 1f37f9d4f3c84757 |
unverified |
MIT (permissive) |
| Unsupervised Deep Embedding for Clustering Analysis |
19 Nov 2015 |
piiswrong/dec/caffe/python/caffe/io.py d47ac8baaac746bc |
unverified |
MIT (permissive) |
| An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition |
21 Jul 2015 |
kurapan/CRNN/utils.py 4dd30f678b132ee5 |
unverified |
MIT (permissive) |
| Conditional Random Fields as Recurrent Neural Networks |
11 Feb 2015 |
torrvision/crfasrnn/python-scripts/crfasrnn_demo.py 31e46c827e832dd6 |
ran · violated contract
fingerprinted |
licence not identified · pointer only |
| Show and Tell: A Neural Image Caption Generator |
17 Nov 2014 |
Cathy-t/HELLO_image/caption/resize.py ddccddb3e642d270 |
unverified |
MIT (permissive) |
| arXiv:openreview_7MlfE2Da2W |
|
snumprlab/scale/experiments/robot/libero/libero_utils.py da6560c251d4e41d |
unverified |
MIT (permissive) |
| arXiv:01705 |
|
yangxy/PASD/pasd/annotator/util.py 50fd68f6989503c6 |
unverified |
Apache-2.0 (permissive) |