{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sliding-puzzles-gym-a-scalable-benchmark-for","title":"Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning","arxiv_id":"2410.14038","date":"2024-10-17","proceeding":null,"authors":["Bryan L. M. de Oliveira","Murilo L. da Luz","Bruno Brandão","Luana G. B. Martins","Telma W. de L. Soares","Luckeciano C. Melo"],"abstract":"Learning effective visual representations is crucial in open-world environments where agents encounter diverse and unstructured observations. This ability enables agents to extract meaningful information from raw sensory inputs, like pixels, which is essential for generalization across different tasks. However, evaluating representation learning separately from policy learning remains a challenge in most reinforcement learning (RL) benchmarks. To address this, we introduce the Sliding Puzzles Gym (SPGym), a benchmark that extends the classic 15-tile puzzle with variable grid sizes and observation spaces, including large real-world image datasets. SPGym allows scaling the representation learning challenge while keeping the latent environment dynamics and algorithmic problem fixed, providing a targeted assessment of agents' ability to form compositional and generalizable state representations. Our experiments with both model-free and model-based RL algorithms, with and without explicit representation learning components, show that as the representation challenge scales, SPGym effectively distinguishes agents based on their capabilities. Moreover, SPGym reaches difficulty levels where no tested algorithm consistently excels, highlighting key challenges and opportunities for advancing representation learning for decision-making research.","url_abs":"https://arxiv.org/abs/2410.14038v2","url_pdf":"https://arxiv.org/pdf/2410.14038v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sliding-puzzles-gym-a-scalable-benchmark-for","repo_url":"https://github.com/bryanoliveira/sliding-puzzles-gym","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.14038","atlas_url":"https://app.syntology.ai/?focus=2410.14038","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.14038"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bryanoliveira/sliding-puzzles-gym","reach":null}],"summary":{"ran_fixture":1,"ran_honours":1,"ran_violates":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"05491384379209b2","entry":"count_inversions","repo":"bryanoliveira/sliding-puzzles-gym","repo_kind":"official","path":"sliding_puzzles/env.py","file_url":"https://github.com/bryanoliveira/sliding-puzzles-gym/blob/HEAD/sliding_puzzles/env.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"05491384379209b2"}},{"code_sha256_prefix":"841b1fd6f21ebe28","entry":"inverse_action","repo":"bryanoliveira/sliding-puzzles-gym","repo_kind":"official","path":"sliding_puzzles/env.py","file_url":"https://github.com/bryanoliveira/sliding-puzzles-gym/blob/HEAD/sliding_puzzles/env.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"841b1fd6f21ebe28"}},{"code_sha256_prefix":"f087e3d610ad7e81","entry":"is_solvable","repo":"bryanoliveira/sliding-puzzles-gym","repo_kind":"official","path":"sliding_puzzles/env.py","file_url":"https://github.com/bryanoliveira/sliding-puzzles-gym/blob/HEAD/sliding_puzzles/env.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f087e3d610ad7e81"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}