Tsinghua's TokenRouter Serving System Delivers 2-64x Decode Speedup for Token-Level LLM Routing
Summary
Tsinghua's NICS-EFC lab released TokenRouter, a serving system for token-level routing across multiple LLMs that reports 2.01x-64.15x higher decoding throughput than existing stacks. The paper attributes gains to a delayed-batching scheduler with analytically derived hyperparameters, plus a request-centric programming model with model-centric execution to tame step desynchronization and batch admission delays. Code is on GitHub (thu-nics/TokenRouter).
Originally reported by huggingface.co
Read the original article →Original headline: Tsinghua's TokenRouter Serving System Delivers 2-64x Decode Speedup for Token-Level LLM Routing