{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-15196","title":"RefDecoder: Enhancing Visual Generation with Conditional Video Decoding","arxiv_id":"2605.15196","date":"2026-05-14","proceeding":null,"authors":["Xiang Fan","Yuheng Wang","Bohan Fang","Zhongzheng Ren","Ranjay Krishna"],"abstract":"Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant loss of detail and inconsistency relative to the input image. To address this, we argue that the decoder requires equal conditioning to preserve structural integrity. We introduce RefDecoder, a reference-conditioned video VAE decoder by injecting high-fidelity reference image signal directly into the decoding process via reference attention. Specifically, a lightweight image encoder maps the reference frame into the detail-rich high-dimensional tokens, which are co-processed with the denoised video latent tokens at each decoder up-sampling stage. We demonstrate consistent improvements across several distinct decoder backbones (e.g., Wan 2.1 and VideoVAE+), achieving up to +2.1dB PSNR over the unconditional baselines on the Inter4K, WebVid, and Large Motion reconstruction benchmarks. Notably, RefDecoder can be directly swapped into existing video generation systems without additional fine-tuning, and we report across-the-board improvements in subject consistency, background consistency, and overall quality scores on the VBench I2V benchmark. Beyond I2V, RefDecoder generalizes well to a wide range of visual generation tasks such as style transfer and video editing refinement.","url_abs":"https://arxiv.org/abs/2605.15196","url_pdf":"https://arxiv.org/pdf/2605.15196","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2605.15196","atlas_url":"https://app.syntology.ai/?focus=2605.15196","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.15196"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/huggingface/diffusers","reach":null}],"summary":{"ran":7,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":8,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6ae68d6c9ac97541","entry":"compute_confidence_aware_loss","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/training_utils.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/training_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6ae68d6c9ac97541"}},{"code_sha256_prefix":"60584419d54b85d5","entry":"compute_snr","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/training_utils.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/training_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"60584419d54b85d5"}},{"code_sha256_prefix":"92d80e4ab339e3ab","entry":"get_constant_schedule","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/optimization.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/optimization.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"92d80e4ab339e3ab"}},{"code_sha256_prefix":"d93f7e15c000f6c3","entry":"get_constant_schedule_with_warmup","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/optimization.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/optimization.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d93f7e15c000f6c3"}},{"code_sha256_prefix":"812bc8a25911527f","entry":"get_piecewise_constant_schedule","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/optimization.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/optimization.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"812bc8a25911527f"}},{"code_sha256_prefix":"6b374abf44ba2349","entry":"is_valid_image","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/image_processor.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/image_processor.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6b374abf44ba2349"}},{"code_sha256_prefix":"a0fb10d3d067e549","entry":"is_valid_image_imagelist","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/image_processor.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/image_processor.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a0fb10d3d067e549"}},{"code_sha256_prefix":"8b349cddb50e4322","entry":"resolve_interpolation_mode","repo":"huggingface/diffusers","repo_kind":"found_in_text","path":"src/diffusers/training_utils.py","file_url":"https://github.com/huggingface/diffusers/blob/HEAD/src/diffusers/training_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8b349cddb50e4322"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9623,"papers_checked":6885},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}