Browse State-of-the-Art › Vision-Language-Action
Vision-Language-Action
49 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 49 papers with code (157 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Jun 2024 3 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedLarge policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from…
-
14 Jul 2025 1 repository listedVision Language Action (VLA) models represent a transformative shift in robotics, with the aim of unifying visual perception, natural language understanding, and embodied control within a single learning framework.
-
7 Jul 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)In this work, we propose VOTE, an efficient and general framework for the optimization and acceleration of VLA models.
-
6 Jul 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation.
-
30 Jun 2025 1 repository listedThe rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single…
-
26 Jun 2025 1 repository listed Syntology ran 0 of 4 samples · 4 unverified · 4 pointer-only (licence)We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation.
-
Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends26 Jun 2025 1 repository listedVLA model post-training aims to address the challenge of improving an embodiment's ability to interact with the environment for the given tasks, analogous to the process of humans motor skills acquisition.
-
16 Jun 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities.
-
16 Jun 2025 1 repository listedThe rapid advancement of generative models has enabled modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specific domains.
-
9 Jun 2025 1 repository listedOur method formulates gesture prediction as a structured sequence denoising task, conditioned on multimodal inputs including endoscopic video, surgical intent language, and a privacy-aware embedding of surgeon identity…
-
9 Jun 2025 1 repository listedTo further reduce the memory footprint of the vision encoder, we propose the distillation-aware training strategy that compresses the full-precision encoder to 1.
-
9 Jun 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedModern AI systems, especially those interacting with the physical world, increasingly require real-time performance.
-
3 Jun 2025 1 repository listedThis differs significantly from LLM jailbreaking literature, as attacks in the real world do not have to be semantically linked to notions of harm.
-
2 Jun 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics.
-
29 May 2025 1 repository listedVision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks.
-
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge28 May 2025 1 repository listed Syntology ran 4 of 17 samples · 13 unverifiedWe argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) Open-world embodied reasoning - the VLA should inherit the knowledge from VLM, i.
-
24 May 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedRecent high-capacity vision-language-action (VLA) models have demonstrated impressive performance on a range of robotic manipulation tasks by imitating human demonstrations.
-
22 May 2025 1 repository listedEmbodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application.
-
21 May 2025 1 repository listedThe generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings.
-
18 May 2025 1 repository listedVision-Language-Action (VLA) models have recently advanced robotic manipulation by translating natural-language instructions and image information into sequential control actions.
-
13 May 2025 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedTo address these limitations, we propose FSD (From Seeing to Doing), a novel vision-language model that generates intermediate representations through spatial relationship reasoning, providing fine-grained guidance for…
-
9 May 2025 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedLearned from internet-scale videos, the generalist policy can be deployed to various robots through efficient latent action decoding.
-
8 May 2025 1 repository listedVision-language-action (VLA) models represent an important step toward general-purpose robotic systems by integrating visual perception, language understanding, and action execution.
-
6 May 2025 1 repository listed Syntology ran 0 of 12 samples · 12 unverifiedDual-system VLA (Vision-Language-Action) architectures have become a hot topic in embodied intelligence research, but there is a lack of sufficient open-source work for further performance analysis and optimization.
-
14 Apr 2025 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedExisting efforts in building Graphical User Interface (GUI) agents largely rely on the training paradigm of supervised fine-tuning on Large Vision-Language Models (LVLMs).
-
30 Mar 2025 1 repository listed Syntology ran 2 of 9 samples · 7 unverifiedWe present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving.
-
25 Mar 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedWhile recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or…
-
12 Mar 2025 1 repository listedRecent advances in Vision-Language-Action models (VLAs) have expanded the capabilities of embodied intelligence.
-
10 Mar 2025 1 repository listed Syntology ran 7 of 15 samples · 8 unverifiedVision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning critical for real-world interaction.
-
27 Feb 2025 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedIn real-world evaluations, our fine-tuning recipe enables OpenVLA to successfully execute dexterous, high-frequency control tasks on a bimanual ALOHA robot and outperform other VLAs (π₀ and RDT-1B) fine-tuned using…
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections