| WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering added by Syntology |
2026-04 (from id) |
zhuyjan/WikiSeeker/utils/utils.py d862607dc4bbee23 |
unverified |
Apache-2.0 (permissive) |
| EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL added by Syntology |
2026-02 (from id) |
LunjunZhang/ema-pg/math/rllm/misc.py 8e21ff6e15adbc8f |
unverified |
MIT (permissive) |
| BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning added by Syntology |
2026-02 (from id) |
YubaoZhao/BranPO/rllm/misc.py 8e21ff6e15adbc8f |
unverified |
Apache-2.0 (permissive) |
| Towards Visual Grounding: A Survey |
28 Dec 2024 |
linhuixiao/clip-vg/pseudo_label_generation_module/pseudo_caption_label_generation/CLIP_prefix_caption/parse_conceptual.py 0a286fb23ba005b9 |
unverified |
Apache-2.0 (permissive) |
| AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving |
19 Dec 2024 |
taco-group/autotrust/Dolphins/inference.py 4aa119881c0ba0a5 |
unverified |
Apache-2.0 (permissive) |
| AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving |
19 Dec 2024 |
taco-group/autotrust/Dolphins/inference_image.py 2baad8535868b50e |
unverified |
Apache-2.0 (permissive) |
| ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models |
9 Dec 2024 |
jieyuz2/provision/provision/annotation/utils.py 4354809003526c37 |
unverified |
Apache-2.0 (permissive) |
| Robot Policy Learning with Temporal Optimal Transport Reward |
29 Oct 2024 |
fuyw/TemporalOT/utils/train_utils.py d7256470f81ee1fb |
unverified |
no licence file found · pointer only |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/data_processing/utils.py 1f804baed33a101d |
unverified |
no licence file found · pointer only |
| WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines |
16 Oct 2024 |
worldcuisines/worldcuisines/evaluation/gemini.py 59daa7bc4cd39ffa |
unverified |
Apache-2.0 (permissive) |
| MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling |
14 Oct 2024 |
tmelyralab/musetalk/musetalk/utils/blending.py ee234e9bbf144430 |
ran
|
licence not identified · pointer only |
| TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation |
19 Sep 2024 |
liyaxuanliyaxuan/TinyVLA/eval_real_franka.py 05b9e2bfc9908f24 |
unverified |
MIT (permissive) |
| CogVLM2: Visual Language Models for Image and Video Understanding |
29 Aug 2024 |
thudm/glm-4/inference/trans_web_vision_demo.py 3ff3af9bca0ae239 |
ran
|
Apache-2.0 (permissive) |
| Fast Context-Based Low-Light Image Enhancement via Neural Implicit Representations |
17 Jul 2024 |
ctom2/colie/utils.py ff11c5f2ecf6311f |
ran
|
Apache-2.0 (permissive) |
| ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools |
18 Jun 2024 |
thudm/chatglm4/inference/trans_web_vision_demo.py 3ff3af9bca0ae239 |
ran
|
Apache-2.0 (permissive) |
| ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO |
17 Jun 2024 |
snumprlab/SRT/data_processing/utils.py 1f804baed33a101d |
unverified |
no licence file found · pointer only |
| The Platonic Representation Hypothesis |
13 May 2024 |
minyoungg/platonic-rep/data.py 56bb81d6b982f638 |
ran
|
no licence file found · pointer only |
| Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond |
15 Mar 2024 |
weitianxin/UniMP/UniMP/pipeline/eval/benchmark_otter.py 3217afe4fa565c58 |
unverified |
no licence file found · pointer only |
| Visual Style Prompting with Swapping Self-Attention |
20 Feb 2024 |
naver-ai/Visual-Style-Prompting/visualize_attention_src/utils.py 93d7815126d56733 |
ran
|
Apache-2.0 (permissive) |
| Separating common from salient patterns with Contrastive Representation Learning |
19 Feb 2024 |
neurospin-projects/2024_rlouiset_sep_clr/celeba_accessories/create_dataset.py dca13cb42b45478e |
ran · honoured contract
|
no licence file found · pointer only |
| The Instinctive Bias: Spurious Images lead to Illusion in MLLMs |
6 Feb 2024 |
MasaiahHan/CorrelationQA/eval_llava.py 60c0cbaaa74fa420 |
ran
|
no licence file found · pointer only |
| Frequency Spectrum is More Effective for Multimodal Representation and Fusion: A Multimodal Spectrum Rumor Detector |
18 Dec 2023 |
dm4m/fsru/data_loader.py 5f96ab7604c7e01d |
ran
|
no licence file found · pointer only |
| Dolphins: Multimodal Language Model for Driving |
1 Dec 2023 |
safolab-wisc/dolphins/inference.py 4aa119881c0ba0a5 |
unverified |
MIT (permissive) |
| OtterHD: A High-Resolution Multi-modality Model |
7 Nov 2023 |
luodian/otter/pipeline/demos/interactive/otter_image.py fe65bfdd78032556 |
ran
|
MIT (permissive) |
| OtterHD: A High-Resolution Multi-modality Model |
7 Nov 2023 |
luodian/otter/pipeline/demos/interactive/otter_video.py bdd50d5971d6d069 |
ran
|
MIT (permissive) |
| OtterHD: A High-Resolution Multi-modality Model |
7 Nov 2023 |
luodian/otter/pipeline/demos/interactive/otter_image_incontext.py 8dab437d86150725 |
unverified |
MIT (permissive) |
| Evaluating Object Hallucination in Large Vision-Language Models |
17 May 2023 |
rucaibox/pope/utils.py d812ee50bd3ae6d8 |
unverified |
MIT (permissive) |
| Revisiting Implicit Neural Representations in Low-Level Vision |
20 Apr 2023 |
WenTXuL/LINR/train-denoising.py fd7ec7b182289996 |
ran · honoured contract
|
no licence file found · pointer only |
| Exploring Discrete Diffusion Models for Image Captioning |
21 Nov 2022 |
buxiangzhiren/ddcap/parse_conceptual.py 0a286fb23ba005b9 |
unverified |
MIT (permissive) |
| Generative Principal Component Analysis |
18 Mar 2022 |
liuzq09/GenerativePCA/src/celebA_estimators.py 5be799e3ef989d15 |
ran · honoured contract
|
no licence file found · pointer only |
| Graph-Augmented Normalizing Flows for Anomaly Detection of Multiple Time Series |
16 Feb 2022 |
khalooei/ALOCC-CVPR2018/utils.py a0b95d6b4e00d081 |
unverified |
MIT (permissive) |
| ClipCap: CLIP Prefix for Image Captioning |
18 Nov 2021 |
rmokady/clip_prefix_caption/parse_conceptual.py 0a286fb23ba005b9 |
unverified |
MIT (permissive) |
| E-RAFT: Dense Optical Flow from Event Cameras |
24 Aug 2021 |
uzh-rpg/e-raft/loader/utils.py ca480d2713282e55 |
unverified |
MIT (permissive) |
| RetinaFace: Single-stage Dense Face Localisation in the Wild |
2 May 2019 |
qiaoxiu/facenetRetinaFace/face/infer2.py 006737632500ead2 |
unverified |
MIT (permissive) |
| COCO-GAN: Generation by Parts via Conditional Coordinating |
30 Mar 2019 |
hubert0527/COCO-GAN/img_utils.py 068cd5ab0df5f40c |
unverified |
MIT (permissive) |
| Contrastive Variational Autoencoder Enhances Salient Features |
12 Feb 2019 |
abidlabs/contrastive_vae/helper.py 7e1a84dc738a0cf2 |
unverified |
MIT (permissive) |
| Task2Vec: Task Embedding for Meta-Learning |
10 Feb 2019 |
awslabs/aws-cv-task2vec/plot_distance_cub_inat.py 052a7c3d285c564d |
unverified |
Apache-2.0 (permissive) |
| A Style-Based Generator Architecture for Generative Adversarial Networks |
12 Dec 2018 |
yan-roo/FakeFace/helper.py 7e1a84dc738a0cf2 |
unverified |
MIT (permissive) |
| Robustness of Conditional GANs to Noisy Labels |
8 Nov 2018 |
POLane16/Robust-Conditional-GAN/mnist/utils.py a0b95d6b4e00d081 |
unverified |
MIT (permissive) |
| Robustness of Conditional GANs to Noisy Labels |
8 Nov 2018 |
POLane16/Robust-Conditional-GAN/cifar10/common/misc.py c69f8aa044146a22 |
unverified |
MIT (permissive) |
| MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices |
20 Apr 2018 |
foamliu/MobileFaceNet/custom_eval.py 0a6e1a8eb8473974 |
unverified |
Apache-2.0 (permissive) |
| Solving Linear Inverse Problems Using GAN Priors: An Algorithm with Provable Guarantees |
23 Feb 2018 |
shahviraj/pgdgan/src/dcgan_utils.py 0639b692497f7a55 |
unverified |
MIT (permissive) |
| Query-Efficient Black-box Adversarial Examples (superceded) |
19 Dec 2017 |
identical code first harvested elsewhere 197a1c1869bf9116 |
unverified |
licence of this copy not recorded |
| StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks |
19 Oct 2017 |
hanzhanggit/StackGAN/misc/utils.py 21bf8f7cea6decf5 |
unverified |
MIT (permissive) |
| Progressive Growing of GANs for Improved Quality, Stability, and Variation |
27 Oct 2017 |
alexeyhorkin/ProGAN-PyTorch/samples/latent_space_interpolation.py 3b64be721c117eb3 |
unverified |
MIT (permissive) |
| Towards Deep Learning Models Resistant to Adversarial Attacks |
19 Jun 2017 |
identical code first harvested elsewhere 197a1c1869bf9116 |
unverified |
licence of this copy not recorded |
| Ensemble Adversarial Training: Attacks and Defenses |
19 May 2017 |
andrewilyas/ens-adv-train-attack/pi-nes.py 197a1c1869bf9116 |
unverified |
no licence file found · pointer only |
| Compressed Sensing using Generative Models |
9 Mar 2017 |
AshishBora/csgm/src/dcgan_utils.py 0639b692497f7a55 |
unverified |
MIT (permissive) |
| Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks |
17 Jan 2017 |
LMescheder/AdversarialVariationalBayes/avb/utils.py de3420581345df2c |
unverified |
MIT (permissive) |
| Finding Tiny Faces |
13 Dec 2016 |
gpspelle/Crowd-Counting/find_people.py 6126b300e57ac811 |
unverified |
MIT (permissive) |
| StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks |
10 Dec 2016 |
rafiahmed40/stack-adverserial-network/misc/utils.py 21bf8f7cea6decf5 |
unverified |
MIT (permissive) |
| Is the deconvolution layer the same as a convolutional layer? |
22 Sep 2016 |
atriumlts/subpixel/utils.py a270ba2f61b9c333 |
unverified |
MIT (permissive) |
| Semantic Image Inpainting with Deep Generative Models |
26 Jul 2016 |
InternityFoundation/SemanticInpainting/image_helpers.py f148395f10b93c68 |
unverified |
no licence file found · pointer only |
| Data-dependent Initializations of Convolutional Neural Networks |
21 Nov 2015 |
cdoersch/deepcontext/utils.py 392fce2a3c3c8a86 |
unverified |
MIT (permissive) |
| Deeply-Supervised Nets |
18 Sep 2014 |
ellisdg/3DUnetCNN/unet3d/utils/image.py 50a36c3ef53a39bf |
unverified |
MIT (permissive) |
| arXiv:2025.emnlp-main.534 |
|
RUCAIBox/POPE/utils.py d812ee50bd3ae6d8 |
unverified |
MIT (permissive) |