Papers › RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling...

RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

8 Mar 2025arXiv:2503.10657archive 2025-07-28

Zhongzhan Huang, Guoming Ling, Yupei Lin, Yandong Chen, Shanshan Zhong, Hefeng Wu, Liang Lin

Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable router can significantly enhance the performance of this paradigm as the number of candidates increases. This improvement can even surpass the performance of the best single model in the pool and many existing strong LLMs, confirming it a highly promising paradigm. However, the lack of comprehensive and open-source benchmarks for Routing LLMs has hindered the development of routers. In this paper, we introduce RouterEval, a benchmark tailored for router research, which includes over 200,000,000 performance records for 12 popular LLM evaluations across various areas such as commonsense reasoning, semantic understanding, etc., based on over 8,500 various LLMs. Using RouterEval, extensive evaluations of existing Routing LLM methods reveal that most still have significant room for improvement. See https://github.com/MilkThink-Lab/RouterEval for all data, code and tutorial.

PaperPDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2503.10657")

Code

Syntology Ran 0 of 13 code samples harvested from 1 repository linked to this paper; 13 have no recorded run.

By repository: official repository: 13 samples from 1 repository, 0 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

milkthink-lab/routereval officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

13 samples harvested; 0 ran; 0 honoured the contract we drafted; 13 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

13unverified

Licence: 0 of the 13 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from milkthink-lab/routereval. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

a3m_ensemble_route milkthink-lab/routereval/router/A3M-router/a3m_router.py official repository unverified MIT (permissive) · deff5e9b4c4a78eb · report
constant milkthink-lab/routereval/router/MLPR_LinearR/tools.py official repository unverified MIT (permissive) · 2713e130185a23e5 · report
cosine_lr milkthink-lab/routereval/router/MLPR_LinearR/tools.py official repository unverified MIT (permissive) · 7529a26792fd6da5 · report
create_embeds milkthink-lab/routereval/utils.py official repository unverified MIT (permissive) · 8e6400b1962406a9 · report
create_prompts milkthink-lab/routereval/utils.py official repository unverified MIT (permissive) · 6bc2f9a20da48d95 · report
dataset_perpare milkthink-lab/routereval/router/RoBERTa-MLC/roberta_MLC.py official repository unverified MIT (permissive) · fd82a32520c85a92 · report
mix_router milkthink-lab/routereval/router/R_o/r_o_router.py official repository unverified MIT (permissive) · 50768d80fbec340b · report
oracle_router milkthink-lab/routereval/get_router_dataset.py official repository unverified MIT (permissive) · 9a63a1d713d235f6 · report
predict_proba milkthink-lab/routereval/router/C-RoBERTa-cluster/cluster.py official repository unverified MIT (permissive) · 1f7e67d5b3462795 · report
prepare_data milkthink-lab/routereval/utils.py official repository unverified MIT (permissive) · 7c51505a5adb6750 · report
process_label milkthink-lab/routereval/get_router_dataset.py official repository unverified MIT (permissive) · 80695e8696cdbc52 · report
split_data milkthink-lab/routereval/get_router_dataset.py official repository unverified MIT (permissive) · d704eae539e693d2 · report
step_lr milkthink-lab/routereval/router/MLPR_LinearR/tools.py official repository unverified MIT (permissive) · 3ea5687c1026fe43 · report

Tasks

Instruction FollowingMathematical Reasoning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections