{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/agentohana-design-unified-data-and-training","title":"AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning","arxiv_id":"2402.15506","date":"2024-02-23","proceeding":null,"authors":["JianGuo Zhang","Tian Lan","Rithesh Murthy","Zhiwei Liu","Weiran Yao","Ming Zhu","Juntao Tan","Thai Hoang","Zuxin Liu","Liangwei Yang","Yihao Feng","Shirley Kokane","Tulika Awalgaonkar","Juan Carlos Niebles","Silvio Savarese","Shelby Heinecke","Huan Wang","Caiming Xiong"],"abstract":"Autonomous agents powered by large language models (LLMs) have garnered significant research attention. However, fully harnessing the potential of LLMs for agent-based tasks presents inherent challenges due to the heterogeneous nature of diverse data sources featuring multi-turn trajectories. In this paper, we introduce \\textbf{AgentOhana} as a comprehensive solution to address these challenges. \\textit{AgentOhana} aggregates agent trajectories from distinct environments, spanning a wide array of scenarios. It meticulously standardizes and unifies these trajectories into a consistent format, streamlining the creation of a generic data loader optimized for agent training. Leveraging the data unification, our training pipeline maintains equilibrium across different data sources and preserves independent randomness across devices during dataset partitioning and model training. Additionally, we present \\textbf{xLAM-v0.1}, a large action model tailored for AI agents, which demonstrates exceptional performance across various benchmarks. Begin the exploration at \\url{https://github.com/SalesforceAIResearch/xLAM}.","url_abs":"https://arxiv.org/abs/2402.15506v4","url_pdf":"https://arxiv.org/pdf/2402.15506v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"agentohana-design-unified-data-and-training","repo_url":"https://github.com/SalesforceAIResearch/xLAM/tree/main/xLAM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"agentohana-design-unified-data-and-training","repo_url":"https://github.com/SalesforceAIResearch/xLAM","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.15506","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.15506"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/SalesforceAIResearch/xLAM","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/SalesforceAIResearch/xLAM/tree/main/xLAM","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"64697db937136197","entry":"dispatch_api_requests","repo":"SalesforceAIResearch/xLAM","repo_kind":"official","path":"actionstudio/src/data_pipeline/criticLAM/trajectory_critic.py","file_url":"https://github.com/SalesforceAIResearch/xLAM/blob/HEAD/actionstudio/src/data_pipeline/criticLAM/trajectory_critic.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"64697db937136197"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}