{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-multi-task-by-active-sampling","title":"Learning to Multi-Task by Active Sampling","arxiv_id":"1702.06053","date":"2017-02-20","proceeding":"ICLR 2018 1","authors":["Sahil Sharma","Ashutosh Jha","Parikshit Hegde","Balaraman Ravindran"],"abstract":"One of the long-standing challenges in Artificial Intelligence for learning\ngoal-directed behavior is to build a single agent which can solve multiple\ntasks. Recent progress in multi-task learning for goal-directed sequential\nproblems has been in the form of distillation based learning wherein a student\nnetwork learns from multiple task-specific expert networks by mimicking the\ntask-specific policies of the expert networks. While such approaches offer a\npromising solution to the multi-task learning problem, they require supervision\nfrom large expert networks which require extensive data and computation time\nfor training. In this work, we propose an efficient multi-task learning\nframework which solves multiple goal-directed tasks in an on-line setup without\nthe need for expert supervision. Our work uses active learning principles to\nachieve multi-task learning by sampling the harder tasks more than the easier\nones. We propose three distinct models under our active sampling framework. An\nadaptive method with extremely competitive multi-tasking performance. A\nUCB-based meta-learner which casts the problem of picking the next task to\ntrain on as a multi-armed bandit problem. A meta-learning method that casts the\nnext-task picking problem as a full Reinforcement Learning problem and uses\nactor critic methods for optimizing the multi-tasking performance directly. We\ndemonstrate results in the Atari 2600 domain on seven multi-tasking instances:\nthree 6-task instances, one 8-task instance, two 12-task instances and one\n21-task instance.","url_abs":"http://arxiv.org/abs/1702.06053v4","url_pdf":"http://arxiv.org/pdf/1702.06053v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-multi-task-by-active-sampling","repo_url":"https://github.com/andris955/diplomaterv","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1702.06053","atlas_url":"https://app.syntology.ai/?focus=1702.06053","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}