{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-23057","title":"RequestRouter: Request-Boundary Routing for Efficient Single-GPU LLM Inference","arxiv_id":"2605.23057","date":"2026-05-21","proceeding":null,"authors":["Aman Sunesh","Ali Alshehhi","Hivansh Dhakne"],"abstract":"RequestRouter is a lightweight request-boundary controller for reducing the latency and energy cost of single-GPU large language model inference. Rather than serving all requests with one static configuration, the system uses cheap request-level features to select one fixed inference mode per request, including FP16, quantized inference, speculative decoding, prefix caching, continuous batching, and hybrid modes such as GPTQ plus prefix caching and INT8 plus continuous batching. We evaluate RequestRouter using an 8B instruction-tuned language model served through vLLM on NVIDIA A100 GPUs. Across the full-scale A100 evaluation--26,500 fixed-mode evaluations followed by 3,500 online-controller evaluations, for 30,000 measured inference executions in total--the controller achieves a 2.10x mean latency speedup over FP16 and a 0.48x energy ratio on deployment-style workloads. A smaller matched evaluation with repeated measurements provides a controlled statistical check of this result: RequestRouter retains a 1.93x latency speedup (95% CI: 1.88--1.98x) and a 0.523 energy ratio (95% CI: 0.506--0.540), showing that the gains persist under a more tightly controlled protocol. On a separate expanded automatic benchmark evaluation, the routed policy retains 99.6% of FP16 macro accuracy. A 100,000-call CPU microbenchmark measures only 0.00475 ms mean routing overhead (0.00532 ms p99). Thus, simple request-aware routing can recover substantial serving efficiency without retraining or modifying the underlying LLM.","url_abs":"https://arxiv.org/abs/2605.23057","url_pdf":"https://arxiv.org/pdf/2605.23057","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2605.23057","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.23057"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM","reach":null}],"summary":{"ran":3,"ran_draft_wrong":2,"ran_honours":1},"by_repo_kind":{"found_in_text":{"samples":6,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9f6818594bc4bf29","entry":"ControllerClassification","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9f6818594bc4bf29"}},{"code_sha256_prefix":"ee216657a7d9152f","entry":"ControllerDecision","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ee216657a7d9152f"}},{"code_sha256_prefix":"6fa9dd1405869b1d","entry":"RequestFeatures","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6fa9dd1405869b1d"}},{"code_sha256_prefix":"694bd06ac048b6e8","entry":"classify_request","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"694bd06ac048b6e8"}},{"code_sha256_prefix":"ea48819b99976ef9","entry":"estimate_prefill_share_pct","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ea48819b99976ef9"}},{"code_sha256_prefix":"71c2d9f55d5cc599","entry":"route_request","repo":"ModeSwitch-LLM/ModeSwitch-LLM","repo_kind":"found_in_text","path":"controller/router.py","file_url":"https://github.com/ModeSwitch-LLM/ModeSwitch-LLM/blob/HEAD/controller/router.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"71c2d9f55d5cc599"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,795 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6795},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}