{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sensitive-and-scalable-online-evaluation-with","title":"Sensitive and Scalable Online Evaluation with Theoretical Guarantees","arxiv_id":"1711.09454","date":"2017-11-26","proceeding":null,"authors":["Oosterhuis Harrie","de Rijke Maarten"],"abstract":"Multileaved comparison methods generalize interleaved comparison methods to\nprovide a scalable approach for comparing ranking systems based on regular user\ninteractions. Such methods enable the increasingly rapid research and\ndevelopment of search engines. However, existing multileaved comparison methods\nthat provide reliable outcomes do so by degrading the user experience during\nevaluation. Conversely, current multileaved comparison methods that maintain\nthe user experience cannot guarantee correctness. Our contribution is two-fold.\nFirst, we propose a theoretical framework for systematically comparing\nmultileaved comparison methods using the notions of considerateness, which\nconcerns maintaining the user experience, and fidelity, which concerns reliable\ncorrect outcomes. Second, we introduce a novel multileaved comparison method,\nPairwise Preference Multileaving (PPM), that performs comparisons based on\ndocument-pair preferences, and prove that it is considerate and has fidelity.\nWe show empirically that, compared to previous multileaved comparison methods,\nPPM is more sensitive to user preferences and scalable with the number of\nrankers being compared.","url_abs":"http://arxiv.org/abs/1711.09454v1","url_pdf":"http://arxiv.org/pdf/1711.09454v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sensitive-and-scalable-online-evaluation-with","repo_url":"https://github.com/HarrieO/PairwisePreferenceMultileave","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}