Papers › Optimizing Attention and Cognitive Control Costs Using Temporally-Layered Architectures
Optimizing Attention and Cognitive Control Costs Using Temporally-Layered Architectures
Devdhar Patel, Terrence Sejnowski, Hava Siegelmann
The current reinforcement learning framework focuses exclusively on performance, often at the expense of efficiency. In contrast, biological control achieves remarkable performance while also optimizing computational energy expenditure and decision frequency. We propose a Decision Bounded Markov Decision Process (DB-MDP), that constrains the number of decisions and computational energy available to agents in reinforcement learning environments. Our experiments demonstrate that existing reinforcement learning algorithms struggle within this framework, leading to either failure or suboptimal performance. To address this, we introduce a biologically-inspired, Temporally Layered Architecture (TLA), enabling agents to manage computational costs through two layers with distinct time scales and energy requirements. TLA achieves optimal performance in decision-bounded environments and in continuous control environments, it matches state-of-the-art performance while utilizing a fraction of the compute cost. Compared to current reinforcement learning algorithms that solely prioritize performance, our approach significantly lowers computational energy expenditure while maintaining performance. These findings establish a benchmark and pave the way for future research on energy and time-aware control.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| OpenAI Gym | Ant-v2 | TLA | Action Repetition | .1268 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | Ant-v2 | TLA | Average Decisions | 860.21 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | Ant-v2 | TLA | Mean Reward | 5163.54 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | HalfCheetah-v2 | TLA | Action Repetition | .1805 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | HalfCheetah-v2 | TLA | Average Decisions | 831.42 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | HalfCheetah-v2 | TLA | Mean Reward | 9571.99 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | Hopper-v2 | TLA | Action Repetition | .5722 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | Hopper-v2 | TLA | Average Decisions | 423.91 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | Hopper-v2 | TLA | Mean Reward | 3458.22 | #1 of 2 | Archive leaderboard | report |
| OpenAI Gym | InvertedDoublePendulum-v2 | TLA | Action Repetition | .7522 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | InvertedDoublePendulum-v2 | TLA | Average Decisions | 247.76 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | InvertedDoublePendulum-v2 | TLA | Mean Reward | 9356.67 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | InvertedPendulum-v2 | TLA | Action Repetition | .8882 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | InvertedPendulum-v2 | TLA | Average Decisions | 111.79 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | InvertedPendulum-v2 | TLA | Mean Reward | 1000 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | MountainCarContinuous-v0 | TLA | Action Repetition | .914 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | MountainCarContinuous-v0 | TLA | Average Decisions | 10.6 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | MountainCarContinuous-v0 | TLA | Mean Reward | 93.88 | #1 of 1 | Archive leaderboard | report |
| OpenAI Gym | Pendulum-v1 | TLA | Action Repetition | .7032 | #2 of 2 | Archive leaderboard | report |
| OpenAI Gym | Pendulum-v1 | TLA | Average Decisions | 62.31 | #2 of 2 | Archive leaderboard | report |
| OpenAI Gym | Pendulum-v1 | TLA | Mean Reward | -154.92 | #2 of 2 | Archive leaderboard | report |
| OpenAI Gym | Walker2d-v2 | TLA | Action Repetition | .4745 | #2 of 2 | Archive leaderboard | report |
| OpenAI Gym | Walker2d-v2 | TLA | Average Decisions | 513.12 | #2 of 2 | Archive leaderboard | report |
| OpenAI Gym | Walker2d-v2 | TLA | Mean Reward | 3878.41 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: TLA
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections