{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2604-00886","title":"PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding","arxiv_id":"2604.00886","date":"2026-04-01","proceeding":null,"authors":["Nan Wang","Zhiwei Jin","Chen Chen","Haonan Lu"],"abstract":"Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine-grained text and small UI elements demand high-resolution inputs that produce tens of thousands of visual tokens. We observe that this cost is largely wasteful -- across document and GUI benchmarks, only 22--71\\% of image patches are pixel-unique, the rest being exact duplicates of another patch in the same image. We propose \\textbf{PixelPrune}, which exploits this pixel-level redundancy through predictive-coding-based compression, pruning redundant patches \\emph{before} the Vision Transformer (ViT) encoder. Because it operates in pixel space prior to any neural computation, PixelPrune accelerates both the ViT encoder and the downstream LLM, covering the full inference pipeline. The method is training-free, requires no learnable parameters, and supports pixel-lossless compression ($τ{=}0$) as well as controlled lossy compression ($τ{>}0$). Experiments across three model scales and document and GUI benchmarks show that PixelPrune maintains competitive task accuracy while delivering up to 4.2$\\times$ inference speedup and 1.9$\\times$ training acceleration. Code is available at https://github.com/OPPO-Mente-Lab/PixelPrune.","url_abs":"https://arxiv.org/abs/2604.00886","url_pdf":"https://arxiv.org/pdf/2604.00886","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2604.00886","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2604.00886"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/OPPO-Mente-Lab/PixelPrune","reach":null}],"summary":{"ran":1,"ran_fixture":1,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"05c4ecb2ffc70c11","entry":"Pred2DSelector","repo":"OPPO-Mente-Lab/PixelPrune","repo_kind":"found_in_text","path":"pixelprune/methods/pred_2d.py","file_url":"https://github.com/OPPO-Mente-Lab/PixelPrune/blob/HEAD/pixelprune/methods/pred_2d.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"05c4ecb2ffc70c11"}},{"code_sha256_prefix":"bf2671e3645db8be","entry":"_sim2d","repo":"OPPO-Mente-Lab/PixelPrune","repo_kind":"found_in_text","path":"pixelprune/methods/pred_2d.py","file_url":"https://github.com/OPPO-Mente-Lab/PixelPrune/blob/HEAD/pixelprune/methods/pred_2d.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"bf2671e3645db8be"}},{"code_sha256_prefix":"3f6b36ae0bead8fb","entry":"BasePatchSelector","repo":"OPPO-Mente-Lab/PixelPrune","repo_kind":"found_in_text","path":"pixelprune/methods/pred_2d.py","file_url":"https://github.com/OPPO-Mente-Lab/PixelPrune/blob/HEAD/pixelprune/methods/pred_2d.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3f6b36ae0bead8fb"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}