{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vote-vision-language-action-optimization-with","title":"VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting","arxiv_id":"2507.05116","date":"2025-07-07","proceeding":null,"authors":["Juyi Lin","Amir Taherin","Arash Akbari","Arman Akbari","Lei Lu","Guangyu Chen","Taskin Padir","Xiaomeng Yang","Weiwei Chen","Yiqian Li","Xue Lin","David Kaeli","Pu Zhao","Yanzhi Wang"],"abstract":"Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, their generalization remains limited when applied to novel objects or unfamiliar environments that lie outside the training distribution. To address this, many existing approaches integrate additional components such as depth estimation, segmentation, or even diffusion to improve generalization, at the cost of adding significant computation overhead, resulting in low efficiency. This motivates the exploration of efficient action prediction methods, which are independent of additional high-level visual representations or diffusion techniques. In this work, we propose VOTE, an efficient and general framework for the optimization and acceleration of VLA models. In details, we propose a novel tokenizer-free fine-tuning approach for parallel accurate action prediction, which reduces computational overhead and accelerates inference speed. Additionally, we adopt an ensemble voting strategy for the action sampling, which significantly improves model performance and enhances generalization. Experimental results show that our method achieves state-of-the-art performance with 35$\\times$ faster inference and 145 Hz throughput. All the details and codes will be open-sourced.","url_abs":"https://arxiv.org/abs/2507.05116v1","url_pdf":"https://arxiv.org/pdf/2507.05116v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vote-vision-language-action-optimization-with","repo_url":"https://github.com/LukeLIN-web/VOTE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"vision-language-action","task_name":"Vision-Language-Action"}],"methods":[{"method_slug":"adopt","method_name":"ADOPT"},{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2507.05116","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2507.05116"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/LukeLIN-web/VOTE","reach":{"status":"ok"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"834c6a1b85b73658","entry":"check_identical_files","repo":"LukeLIN-web/VOTE","repo_kind":"official","path":"experiments/robot/openvla_utils.py","file_url":"https://github.com/LukeLIN-web/VOTE/blob/HEAD/experiments/robot/openvla_utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"834c6a1b85b73658"}},{"code_sha256_prefix":"543f1c890d52c6aa","entry":"get_image_resize_size","repo":"LukeLIN-web/VOTE","repo_kind":"official","path":"experiments/robot/robot_utils.py","file_url":"https://github.com/LukeLIN-web/VOTE/blob/HEAD/experiments/robot/robot_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"543f1c890d52c6aa"}},{"code_sha256_prefix":"d82a1a874b55c3a6","entry":"log_gpu_memory","repo":"LukeLIN-web/VOTE","repo_kind":"official","path":"experiments/robot/openvla_utils.py","file_url":"https://github.com/LukeLIN-web/VOTE/blob/HEAD/experiments/robot/openvla_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d82a1a874b55c3a6"}},{"code_sha256_prefix":"2495b7375e1e0d26","entry":"model_is_on_hf_hub","repo":"LukeLIN-web/VOTE","repo_kind":"official","path":"experiments/robot/openvla_utils.py","file_url":"https://github.com/LukeLIN-web/VOTE/blob/HEAD/experiments/robot/openvla_utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2495b7375e1e0d26"}},{"code_sha256_prefix":"d3a4b8eedff60e37","entry":"unpack_tuple","repo":"LukeLIN-web/VOTE","repo_kind":"official","path":"prismatic/models/film_vit_wrapper.py","file_url":"https://github.com/LukeLIN-web/VOTE/blob/HEAD/prismatic/models/film_vit_wrapper.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d3a4b8eedff60e37"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}