| Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM added by Syntology |
2026-09 (from id) |
zxp555/ACFT_MM26/ACFT/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
no licence file found · pointer only |
| SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation added by Syntology |
2026-08 (from id) |
Wind-Flipped/SemPOI-RL/code/utils.py 57a8d6752075da4a |
unverified |
no licence file found · pointer only |
| UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers added by Syntology |
2026-08 (from id) |
chidaksh/spurious_mitigator/src/civil_comments/causal_verify_api.py cf80248c006d1518 |
ran
|
no licence file found · pointer only |
| DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models added by Syntology |
2026-08 (from id) |
Zhong-Chenchen/DIVE/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation added by Syntology |
2026-07 (from id) |
opendatalab/MLLM-DataEngine/LLaVA/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study added by Syntology |
2026-07 (from id) |
otanl/microground/src/micro_ground/data.py 0c88aff7f7d23d35 |
ran
|
MIT (permissive) |
| CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence added by Syntology |
2026-06 (from id) |
DGL-codes/CBD/uld/tofuutil/data_module.py cdd171f20cd6f53d |
ran
|
MIT (permissive) |
| Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets? added by Syntology |
2026-06 (from id) |
lin-zhao-resoLve/D3HR/generation/dit_inversion_save_statistic.py 6fc038f89ff59f28 |
ran · our draft was wrong
|
no licence file found · pointer only |
| CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning added by Syntology |
2026-05 (from id) |
zhyzmath/CurveRL/recipe/sppo/sppo_ray_trainer.py a67c1dea604c8a67 |
unverified |
Apache-2.0 (permissive) |
| DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models added by Syntology |
2026-05 (from id) |
CityU-MLO/DualOptimPlus/dataset/data_module.py cdd171f20cd6f53d |
ran
|
no licence file found · pointer only |
| DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning added by Syntology |
2026-05 (from id) |
poudel-bibek/DeCoR/ppo/ppo_utils.py b42728badf61b2c7 |
ran
|
MIT (permissive) |
| Full-Spectrum Graph Neural Networks: Expressive and Scalable added by Syntology |
2026-05 (from id) |
xwangxshi/FSpecGNN/Hom-Cycle-count/main.count.py 7a01234c5af0c132 |
unverified |
no licence file found · pointer only |
| PRIME: Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies added by Syntology |
2026-05 (from id) |
HySonLab/PRIME/plm_baseline.py 890ccf93ef063709 |
ran
|
MIT (permissive) |
| Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL added by Syntology |
2026-04 (from id) |
XIAO4579/PRISM/moe/dense_model/train_dense_vl_warmup.py d36fc4b002a2590e |
ran
|
MIT (permissive) |
| OFA-Diffusion Compression: Compressing Diffusion Model in One-Shot Manner added by Syntology |
2026-04 (from id) |
atrijhy/OFA-Diffusion_Compression/sd/train_ofa.py ffb5e3f9fdba1308 |
unverified |
no licence file found · pointer only |
| Polynomial Expansion Rank Adaptation: Enhancing Low-Rank Fine-Tuning with High-Order Interactions added by Syntology |
2026-04 (from id) |
zhangwenhao6/PERA/dataset/dataset_hg.py 66cc6da5838e8cf0 |
unverified |
licence not identified · pointer only |
| Polynomial Expansion Rank Adaptation: Enhancing Low-Rank Fine-Tuning with High-Order Interactions added by Syntology |
2026-04 (from id) |
zhangwenhao6/PERA/dataset/dataset_hg_combined.py 44b91c3044f0042b |
unverified |
licence not identified · pointer only |
| Back to Basics: Let Conversational Agents Remember with Just Retrieval and Generation added by Syntology |
2026-04 (from id) |
qingyue2014/Rsum/dataloader.py afea9beadf818cbd |
unverified |
no licence file found · pointer only |
| Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions added by Syntology |
2026-04 (from id) |
utshabkg/multi-vector-reproducibility/experiments/10_finetune_constbert_tot.py 5e70a0d689a286bb |
unverified |
no licence file found · pointer only |
| ARES: Scalable and Practical Gradient Inversion Attack in Federated Learning through Activation Recovery added by Syntology |
2026-03 (from id) |
gaow0007/ATSPrivacy/benchmark/cifar100_attack.py 50b843f4577cb9e7 |
unverified |
MIT (permissive) |
| Reforming the Mechanism: Editing Reasoning Patterns in LLMs with Circuit Reshaping added by Syntology |
2026-03 (from id) |
LzyFischer/REdit/src/get_dataset_math.py 30bcd3b2a291a41d |
unverified |
no licence file found · pointer only |
| CRISP: Compressed Reasoning via Iterative Self-Policy Distillation added by Syntology |
2026-03 (from id) |
HJSang/OPSD_Reasoning_Compression/workspace/src/self_distill_hybrid/sd_dataset.py e100d95c89946352 |
unverified |
no licence file found · pointer only |
| SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety added by Syntology |
2026-03 (from id) |
hiyouga/EasyR1/verl/utils/dataset.py 6dabc7804e03b1d0 |
unverified |
Apache-2.0 (permissive) |
| LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards added by Syntology |
2026-03 (from id) |
real-absolute-AI/LongRLVR/recipe/dapo/longrl_reward_manager.py 78baa936d52b4277 |
unverified |
no licence file found · pointer only |
| HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit added by Syntology |
2026-02 (from id) |
EIT-NLP/HiDrop/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| RegionRoute: Regional Style Transfer with Diffusion Model added by Syntology |
2026-02 (from id) |
bwchen05/RegionRoute/training/lora.py dcbeeb19dbf60ca2 |
unverified |
no licence file found · pointer only |
| Agent-Omit: Adaptive Context Omission for Efficient LLM Agents added by Syntology |
2026-02 (from id) |
usail-hkust/Agent-Omit/AgentOmit-RL/verl/utils/agent_dataset/rl_dataset.py ecd75498a47b0cf1 |
unverified |
no licence file found · pointer only |
| TTCS: Test-Time Curriculum Synthesis for Self-Evolving added by Syntology |
2026-01 (from id) |
XMUDeepLIT/TTCS/src/Challenger_dataset.py b8d8d572aa0f7ad9 |
unverified |
no licence file found · pointer only |
| MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning added by Syntology |
2026-01 (from id) |
meituan/MemOCR/recurrent/generation_manager.py f2a7d1d12a2a2a8d |
unverified |
Apache-2.0 (permissive) |
| Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR added by Syntology |
2026-01 (from id) |
Zilai-WANG/Delta-Embedding-Fusion/embed_fusion/data.py fbe00a78b6430e96 |
unverified |
no licence file found · pointer only |
| Contrastive Language-Image Mamba Pretraining added by Syntology |
2026-01 (from id) |
NimrodShabtay/CLIMP/eval_nocaps.py f52bda65e1c22ab1 |
unverified |
no licence file found · pointer only |
| A General Neural Backbone for Mixed-Integer Linear Optimization via Dual Attention added by Syntology |
2026-01 (from id) |
hpx2024/Dual-Attention/02.Element-Level/2_train.py f55c31d782e48ffb |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval added by Syntology |
2026-01 (from id) |
qzp2018/CSMCIR/src/data_utils_csmcir.py b6c1f837e6faaeba |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative added by Syntology |
2026-01 (from id) |
wanglihx/LLaVA-OA/cliptrain_focal.py 0362ac3c9b5bde11 |
unverified |
no licence file found · pointer only |
| WhAM: Towards A Translative Model of Sperm Whale Vocalization added by Syntology |
2025-12 (from id) |
Project-CETI/wham/wham/embedding/train_downstream.py e082a4e27242a338 |
unverified |
MIT (permissive) |
| Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels added by Syntology |
2025-12 (from id) |
OguzBedir/Masked_Symbol_Modeling/core/train_utils.py f847af280cc41e11 |
unverified |
no licence file found · pointer only |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_qwen_m3cot.py 66307ccd5a78a143 |
unverified |
Apache-2.0 (permissive) |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_qwen_mathverse.py 1dcd2f9d386042a6 |
unverified |
Apache-2.0 (permissive) |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_qwen_mmmu.py 2987642d3528ecb0 |
unverified |
Apache-2.0 (permissive) |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_qwen_wemath.py 2a6087e29d42d61d |
unverified |
Apache-2.0 (permissive) |
| Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding added by Syntology |
2025-10 (from id) |
y0ngjaenius/CVPR2024_FLOWERFormer/utils.py 1466d657fc2b6567 |
unverified |
no licence file found · pointer only |
| ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control added by Syntology |
2025-10 (from id) |
tanyuqian/synthetic-private-data/finetune.py 52b43047818c6eae |
unverified |
no licence file found · pointer only |
| QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models added by Syntology |
2025-10 (from id) |
SAI-Lab-NYU/QSVD/fake_quant/eval_utilsdistvizwiz.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models added by Syntology |
2025-10 (from id) |
SAI-Lab-NYU/QSVD/fake_quant/eval_smolvlm_vizwiz.py 07ead1dab90ee7a6 |
unverified |
Apache-2.0 (permissive) |
| SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition added by Syntology |
2025-09 (from id) |
chenshunpeng/SAGE/datasets_ws.py 937b98b0302097b0 |
unverified |
MIT (permissive) |
| Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards added by Syntology |
2025-09 (from id) |
baaivision/CoS/eval/mathvista/evaluate_mathvista.py 9483aaa3e36c23a0 |
unverified |
Apache-2.0 (permissive) |
| Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards added by Syntology |
2025-09 (from id) |
baaivision/CoS/eval/mathvista/evaluate_mathvista_mine.py 53ae416061b7cd9e |
unverified |
Apache-2.0 (permissive) |
| CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning added by Syntology |
2025-08 (from id) |
WNQzhu/CARFT/code/mwp_ReFT/sampling.py 2bb5bd3526f55ecb |
unverified |
no licence file found · pointer only |
| arXiv:2507.19993 |
2025-07 (from id) |
Howardkhh/FROSS/EGTR/pretrain_detr.py cdb6e23f77eb4b45 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.17342 |
2025-07 (from id) |
fudan-zvg/DeMo/src/datamodule/av2_dataset.py cb3df1ace0ca2089 |
unverified |
no licence file found · pointer only |
| CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions |
8 Jul 2025 |
lukahhcm/cultureclip/model_finetune/data.py 041e7ddf4cb7d30f |
unverified |
no licence file found · pointer only |
| LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs |
1 Jul 2025 |
CnFaker/LLaVA-SP/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments |
30 Jun 2025 |
Feliciaxyao/NavMorph/vlnce_baselines/dagger_trainer.py f474c29ed36d8065 |
unverified |
no licence file found · pointer only |
| Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs |
21 Jun 2025 |
aoshuang92/splora/llama2_alpaca_pd_lora.py 3053998047069bec |
unverified |
no licence file found · pointer only |
| UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation |
20 Jun 2025 |
tliby/unifork/unifork/eval/model_vqa_loader.py 9738a62bfa4c292d |
unverified |
no licence file found · pointer only |
| arXiv:2506.17124 |
2025-06 (from id) |
prediction-action-lab/thinking-as-control/data_utils.py 29defc8158c0ead0 |
unverified |
no licence file found · pointer only |
| Watermarking Autoregressive Image Generation |
19 Jun 2025 |
GAIR-NLP/anole/facilitating_image_generation/train_image_head.py abbbb29b34ebbb42 |
unverified |
no licence file found · pointer only |
| Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs |
12 Jun 2025 |
theia-4869/cdpruner/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay |
5 Jun 2025 |
astral-group/data-efficient-llm-rl/rl_training/verl/verl/trainer/ppo/ray_trainer_teacher_replay.py d11b448b5d76e167 |
unverified |
no licence file found · pointer only |
| Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos |
5 Jun 2025 |
AFeng-x/Draw-and-Understand/accessory/eval/infer_and_save.py f5b33c6a8295f129 |
unverified |
Apache-2.0 (permissive) |
| Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency |
2 Jun 2025 |
appletea233/temporal-r1/verl/utils/dataset.py 69132eba50f047ff |
unverified |
Apache-2.0 (permissive) |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/eval/model_vqa_loader_batch.py 5ef2d21cca6281db |
unverified |
Apache-2.0 (permissive) |
| OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation |
26 May 2025 |
PKU-YuanGroup/OpenS2V-Nexus/data_process/step3-1_get_caption.py ea468804bf7f0516 |
unverified |
Apache-2.0 (permissive) |
| OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation |
26 May 2025 |
PKU-YuanGroup/OpenS2V-Nexus/data_process/step4-1_get_tag_local.py 68791d7cdd051179 |
unverified |
Apache-2.0 (permissive) |
| Taming Diffusion for Dataset Distillation with High Representativeness |
23 May 2025 |
lin-zhao-resolve/d3hr/generation/dit_inversion_save_statistic.py 6fc038f89ff59f28 |
ran · our draft was wrong
|
no licence file found · pointer only |
| arXiv:2505.17447 |
2025-05 (from id) |
Cheungki/LeTS/src/verl/workers/reward_manager/lets.py c9608313dd7b30d5 |
unverified |
MIT (permissive) |
| ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay |
22 May 2025 |
dvlab-research/arpo/verl/utils/dataset.py 69132eba50f047ff |
unverified |
Apache-2.0 (permissive) |
| Training-Free Watermarking for Autoregressive Image Generation |
20 May 2025 |
maifoundations/indexmark/Index_encoder/ft2_2loss.py 9780feca2753ce85 |
unverified |
MIT (permissive) |
| s3: You Don't Need That Much Data to Train a Search Agent via RL |
20 May 2025 |
pat-jj/s3/s3/llm_agent/generation_s3.py 90f6865eb340d338 |
unverified |
Apache-2.0 (permissive) |
| MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision |
19 May 2025 |
modalminds/mm-prm/eval/prm/evaluate_k12_prm.py 3426ef0ff637108e |
ran · our draft was wrong
|
MIT (permissive) |
| Visual Planning: Let's Think Only with Images |
16 May 2025 |
yix8/VisualPlanning/train_rl_frozen.py f55c31d782e48ffb |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models |
14 Apr 2025 |
opengvlab/internvl/internvl_chat/eval/mmmu/evaluate_mmmu.py 74fcec52489e905d |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2504.00502 |
2025-04 (from id) |
icip-cas/ShortV/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Sonata: Self-Supervised Learning of Reliable Point Representations |
20 Mar 2025 |
facebookresearch/sonata/sonata/data.py 19a350b1109880a0 |
unverified |
Apache-2.0 (permissive) |
| MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling |
17 Mar 2025 |
hustvl/MaTVLM/tinyllava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM |
17 Mar 2025 |
www-ye/timezero/src/open_r1/sft.py 0ba9682f0587f9f4 |
unverified |
Apache-2.0 (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
MIT (permissive) |
| TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models |
13 Mar 2025 |
shawntan86/tokencarve/TokenCarve/TokenCarve_model_vqa_loader.py 4aa3263a9ac1bd63 |
unverified |
Apache-2.0 (permissive) |
| AdvPaint: Protecting Images from Inpainting Manipulation via Adversarial Attention Disruption |
13 Mar 2025 |
joonsungjeon/advpaint/AdvPaint.py 90fadec0aec03c1b |
ran · honoured contract
|
MIT (permissive) |
| ELECTRA: A Cartesian Network for 3D Charge Density Prediction with Floating Orbitals |
11 Mar 2025 |
Jotels/ELECTRA_2025/models/electra_hotpp/ELECTRA.py f55c31d782e48ffb |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation |
26 Feb 2025 |
3dagentworld/sca3d/SCA3D/dataloaders/data.py eb503fa05caab93e |
unverified |
MIT (permissive) |
| Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment |
24 Feb 2025 |
facico/goat-peft/goat/train_vit.py 9a79ef50064b77e4 |
ran · our draft was wrong
|
MIT (permissive) |
| SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition |
23 Feb 2025 |
Lu-Feng/SelaVPR/datasets_ws.py 937b98b0302097b0 |
unverified |
MIT (permissive) |
| Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More |
17 Feb 2025 |
zichenwen1/dart/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation |
14 Feb 2025 |
dcdmllm/healthgpt/HealthGPT/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Training Sparse Mixture Of Experts Text Embedding Models |
11 Feb 2025 |
nomic-ai/contrastors/src/contrastors/dataset/text_text_loader.py 801ebbc0eea48637 |
unverified |
Apache-2.0 (permissive) |
| TimeKAN: KAN-based Frequency Decomposition Learning Architecture for Long-term Time Series Forecasting |
10 Feb 2025 |
huangst21/timekan/data_provider/uea.py bbdc0c4ddfbaa19b |
ran
|
Apache-2.0 (permissive) |
| UASTHN: Uncertainty-Aware Deep Homography Estimation for UAV Satellite-Thermal Geo-localization |
3 Feb 2025 |
arplaboratory/UASTHN/global_pipeline/datasets_ws.py 247429ce34fc3671 |
ran
|
MIT (permissive) |
| Regularized Langevin Dynamics for Combinatorial Optimization |
1 Feb 2025 |
identical code first harvested elsewhere f55c31d782e48ffb |
ran · honoured contract
fingerprinted |
licence of this copy not recorded |
| Sebra: Debiasing Through Self-Guided Bias Ranking |
30 Jan 2025 |
kadarsh22/Sebra/bar_trainers/sebra.py fbdf12a7504f8ac4 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Multi-concept Model Immunization through Differentiable Model Merging |
19 Dec 2024 |
amberyzheng/MIMA/adaptation/train_custom_diffusion.py 6d82c6864fe791d6 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o |
18 Dec 2024 |
ztangaj/gveval/correlation.py a428a145149517f7 |
unverified |
no licence file found · pointer only |
| ArchesWeather & ArchesWeatherGen: a deterministic and generative model for efficient ML weather forecasting |
17 Dec 2024 |
inria/geoarches/geoarches/main_hydra.py aff28d380df91506 |
unverified |
BSD-3-Clause (permissive) |
| Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models |
17 Dec 2024 |
Savannah120/alignment-handbook-PoFT/generate_preference_scores.py 23bd8e38a433b4b5 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning |
17 Dec 2024 |
ShipingGe/ILCACM/src/util/dataloader.py d3a2efceb5304d00 |
unverified |
MIT (permissive) |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering |
16 Dec 2024 |
bibisbar/LLaVA-Steering/tinyllava/eval/model_vqa_loader.py 68ad5f668d4f2976 |
unverified |
Apache-2.0 (permissive) |
| Causal Graphical Models for Vision-Language Compositional Understanding |
12 Dec 2024 |
aimagelab/COGT/dataset.py bc629eb0de6e6850 |
unverified |
no licence file found · pointer only |
| Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark |
12 Dec 2024 |
yongliang-wu/repurpose/dataset/RepurposeClip.py f1d6f311ec1139dc |
unverified |
MIT (permissive) |
| TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft |
6 Dec 2024 |
teamcraft-bench/teamcraft/llava_teamcraft/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
MIT (permissive) |
| p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay |
5 Dec 2024 |
mcg-nju/p-mod/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation |
4 Dec 2024 |
hqhqaq/patchdpo/inference_dreambooth.py ccf8fb11b2e5dd2c |
unverified |
no licence file found · pointer only |
| A Graph Neural Network Simulation of Dispersed Systems |
2024-12 (from id) |
rfjd/GNS-DispersedSystems/gns/data_loader.py 10d4ddc0d667a1f3 |
unverified |
no licence file found · pointer only |
| Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs |
2 Dec 2024 |
theia-4869/fastervlm/llava/eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification |
1 Dec 2024 |
Osilly/dynamic_llava/llava/dynamic_eval/model_vqa_loader.py 20e4f665698a3d18 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| EDTformer: An Efficient Decoder Transformer for Visual Place Recognition |
1 Dec 2024 |
tong-jin01/edtformer/datasets_ws.py c160834241fc1797 |
unverified |
MIT (permissive) |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/preprocess/beats_feature_extract.py ac11ca70992e72fb |
unverified |
MIT (permissive) |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/preprocess/clip_feature_extract.py 4d8e90858b179363 |
unverified |
MIT (permissive) |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/preprocess/whisper_feature_extract.py 98415dceedab1609 |
unverified |
MIT (permissive) |
| GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks |
28 Nov 2024 |
the-ai-alliance/geo-bench-vlm/eval_geobenchvlm/llava1pt5_cls_single.py ed69126ee53f8ce6 |
unverified |
Apache-2.0 (permissive) |
| GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks |
28 Nov 2024 |
the-ai-alliance/geo-bench-vlm/eval_geobenchvlm/qwen_cls_single.py ca49a9290f4e93c3 |
unverified |
Apache-2.0 (permissive) |