{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/otoworld-towards-learning-to-separate-by","title":"OtoWorld: Towards Learning to Separate by Learning to Move","arxiv_id":"2007.06123","date":"2020-07-12","proceeding":null,"authors":["Omkar Ranadive","Grant Gasser","David Terpay","Prem Seetharaman"],"abstract":"We present OtoWorld, an interactive environment in which agents must learn to listen in order to solve navigational tasks. The purpose of OtoWorld is to facilitate reinforcement learning research in computer audition, where agents must learn to listen to the world around them to navigate. OtoWorld is built on three open source libraries: OpenAI Gym for environment and agent interaction, PyRoomAcoustics for ray-tracing and acoustics simulation, and nussl for training deep computer audition models. OtoWorld is the audio analogue of GridWorld, a simple navigation game. OtoWorld can be easily extended to more complex environments and games. To solve one episode of OtoWorld, an agent must move towards each sounding source in the auditory scene and \"turn it off\". The agent receives no other input than the current sound of the room. The sources are placed randomly within the room and can vary in number. The agent receives a reward for turning off a source. We present preliminary results on the ability of agents to win at OtoWorld. OtoWorld is open-source and available.","url_abs":"https://arxiv.org/abs/2007.06123v1","url_pdf":"https://arxiv.org/pdf/2007.06123v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"otoworld-towards-learning-to-separate-by","repo_url":"https://github.com/pseeth/otoworld","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"audio-source-separation","task_name":"Audio Source Separation"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"openai-gym","task_name":"OpenAI Gym"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.06123","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2007.06123"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pseeth/otoworld","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"45f5fb156f286869","entry":"autoclip","repo":"pseeth/otoworld","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/pseeth/otoworld/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"45f5fb156f286869"}},{"code_sha256_prefix":"552c9ce2c2acfc5d","entry":"compute_ideal_binary_mask","repo":"pseeth/otoworld","repo_kind":"official","path":"src/transforms.py","file_url":"https://github.com/pseeth/otoworld/blob/HEAD/src/transforms.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"552c9ce2c2acfc5d"}},{"code_sha256_prefix":"38c3a53e5808d0e5","entry":"generate_data_from_log","repo":"pseeth/otoworld","repo_kind":"official","path":"src/plot_runs.py","file_url":"https://github.com/pseeth/otoworld/blob/HEAD/src/plot_runs.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"38c3a53e5808d0e5"}},{"code_sha256_prefix":"d474f9a41a26ca78","entry":"ipd_ild_features","repo":"pseeth/otoworld","repo_kind":"official","path":"src/audio_processing.py","file_url":"https://github.com/pseeth/otoworld/blob/HEAD/src/audio_processing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d474f9a41a26ca78"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}