Browse State-of-the-Art › Action Generation
Action Generation
49 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 49 papers with code (111 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
4 Sep 2018 5 repositories listedWe propose to decompose instruction execution to goal prediction and action generation.
-
19 Apr 2025 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecent works have begun exploring reasoning in GUI tasks with encouraging results.
-
4 Feb 2025 2 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 5 pointer-only (licence)We present flow Q-learning (FQL), a simple and performant offline reinforcement learning (RL) method that leverages an expressive flow-matching policy to model arbitrarily complex action distributions in data.
-
19 Apr 2024 2 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedIn this work, we introduce the paradigm of generating web scrapers with LLMs and propose AutoScraper, a two-stage framework that can handle diverse and changing web environments more efficiently.
-
26 May 2018 2 repositories listedInspired by the recent advances in generative models, we introduce a human action generation model in order to generate a consecutive sequence of human motions to formulate novel actions.
-
26 Jun 2025 1 repository listed Syntology ran 0 of 4 samples · 4 unverified · 4 pointer-only (licence)We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation.
-
Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends26 Jun 2025 1 repository listedVLA model post-training aims to address the challenge of improving an embodiment's ability to interact with the environment for the given tasks, analogous to the process of humans motor skills acquisition.
-
16 Jun 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities.
-
5 Jun 2025 1 repository listedFor example, in group chats, online team meetings, or social games, there is no inherent notion of turns; therefore, the decision of when to speak forms a crucial part of the participant's decision making.
-
4 Jun 2025 1 repository listedHowever, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level…
-
4 Jun 2025 1 repository listedTransforming complex actions into discrete skill abstractions has demonstrated strong potential for robotic manipulation.
-
2 Jun 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics.
-
23 May 2025 1 repository listed Syntology ran 1 of 10 samples · 9 unverifiedWe improve agent distillation along two complementary axes: (1) we introduce a prompting method called first-thought prefix to enhance the quality of teacher-generated trajectories; and (2) we propose a self-consistent…
-
15 May 2025 1 repository listedLarge language models (LLMs) have opened new opportunities for automated mobile app exploration, an important and challenging problem that used to suffer from the difficulty of generating meaningful UI interactions.
-
14 May 2025 1 repository listedBy contrast, in action generation, the dimensionality of the target is comparatively small, and only the image condition is high-dimensional.
-
14 Apr 2025 1 repository listedIn visual navigation, previous diffusion-based policies typically generate action sequences by initiating from denoising Gaussian noise.
-
9 Mar 2025 1 repository listed Syntology ran 2 of 14 samples · 12 unverifiedTraditional agentic workflows rely on external prompts to manage interactions with tools and the environment, which limits the autonomy of reasoning models.
-
4 Mar 2025 1 repository listedWe introduce LiteWebAgent, an open-source suite for VLM-based web agent applications.
-
1 Mar 2025 1 repository listedDiffusion models have recently shown significant potential in solving decision-making problems, particularly in generating behavior plans -- also known as diffusion planning.
-
27 Feb 2025 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedIn real-world evaluations, our fine-tuning recipe enables OpenVLA to successfully execute dexterous, high-frequency control tasks on a bimanual ALOHA robot and outperform other VLAs (π₀ and RDT-1B) fine-tuned using…
-
23 Feb 2025 1 repository listedIn this paper, we introduce Action Generation with Plackett-Luce Sampling (AGPS), a novel mechanism for agent decision order optimization.
-
6 Feb 2025 1 repository listedFurthermore, we examine the challenges that limit adapting LLMs in MRS, including mathematical reasoning limitations, hallucination, latency issues, and the need for robust benchmarking systems.
-
13 Dec 2024 1 repository listedAs AI continues to advance, there is a growing demand for systems that go beyond language-based assistance and move toward intelligent agents capable of performing real-world actions.
-
4 Nov 2024 1 repository listedVision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, language comprehension, and action…
-
9 Oct 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Document logical structuring aims to extract the underlying hierarchical structure of documents, which is crucial for document intelligence.
-
2 Sep 2024 1 repository listedOur framework seamlessly unifies affordance learning and action generation with flow matching for robot manipulation.
-
28 Aug 2024 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedLong-horizon decision-making tasks present significant challenges for LLM-based agents due to the need for extensive planning over multiple steps.
-
26 Jul 2024 1 repository listedAs a result, Wonderful Team's performance on real-world semantic and physical planning tasks often exceeds methods that rely on separate vision systems.
-
18 Mar 2024 1 repository listedThe open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated.
-
2 Feb 2024 1 repository listedWe introduce PokeLLMon, the first LLM-embodied agent that achieves human-parity performance in tactical battle games, as demonstrated in Pokemon battles.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections