{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/13062597","title":"Introducing LETOR 4.0 Datasets","arxiv_id":"1306.2597","date":"2013-06-09","proceeding":null,"authors":["Tao Qin","Tie-Yan Liu"],"abstract":"LETOR is a package of benchmark data sets for research on LEarning TO Rank,\nwhich contains standard features, relevance judgments, data partitioning,\nevaluation tools, and several baselines. Version 1.0 was released in April\n2007. Version 2.0 was released in Dec. 2007. Version 3.0 was released in Dec.\n2008. This version, 4.0, was released in July 2009. Very different from\nprevious versions (V3.0 is an update based on V2.0 and V2.0 is an update based\non V1.0), LETOR4.0 is a totally new release. It uses the Gov2 web page\ncollection (~25M pages) and two query sets from Million Query track of TREC\n2007 and TREC 2008. We call the two query sets MQ2007 and MQ2008 for short.\nThere are about 1700 queries in MQ2007 with labeled documents and about 800\nqueries in MQ2008 with labeled documents. If you have any questions or\nsuggestions about the datasets, please kindly email us (letor@microsoft.com).\nOur goal is to make the dataset reliable and useful for the community.","url_abs":"http://arxiv.org/abs/1306.2597v1","url_pdf":"http://arxiv.org/pdf/1306.2597v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"13062597","repo_url":"https://github.com/kiwicom/catboost-cxx","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"13062597","repo_url":"https://github.com/sajari/mlg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"13062597","repo_url":"https://github.com/wildltr/ptranking","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"}],"methods":[],"datasets_introduced":[{"slug":"mq2007","name":"MQ2007","full_name":""},{"slug":"mq2008","name":"MQ2008","full_name":""},{"slug":"mslr-web30k","name":"MSLR WEB30K","full_name":"Microsoft Learning to Rank Datasets-30k"},{"slug":"mslr-web10k","name":"MSLR-WEB10K","full_name":""},{"slug":"mslr-web30k-1","name":"MSLR-WEB30K","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1306.2597","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}