Yandex/Together AI: Training-Free AsyncLLM Runs 8 Parallel Coroutines at 552 tok/s on 9B

Found first: a primary source the press has not covered yet.

Researchers from Yandex, Together AI, HSE University, and Yandex School of Data Analysis have released a paper introducing AsyncLLM, a training-free Python framework that lets any existing LLM run as concurrent inference coroutines with overlapping memory states. Posted to arXiv on 28 September 2026, the work shows that off-the-shelf models can accept new inputs mid-inference, tested across streaming video, system monitoring, and videogame benchmarks, without any fine-tuning.

What the source says

George Yakushev, Denis Mazur, and five co-authors built the framework around Python async/await coroutines that share cache blocks holding attention KV caches and recurrent states, batched across a single GPU using attention manipulation. Tested models include Qwen3.5-9B, Qwen3.8-27B, and Qwen3.6-35B-A3B. On SoccerNet-Caption streaming video, Qwen3.5-9B under AsyncLLM reached TriggerAcc 0.677 against a Mage-VL sequential baseline of 0.555. On DevOps-Gym monitoring (34 tasks), Qwen3.6-35B-A3B achieved 55.88% accuracy using 3,837 forward passes, compared to 61.76% for the sequential baseline at 8,452 passes. Decoding throughput on a single H200 scaled from 106 tokens/sec at 1 coroutine to 552 tokens/sec at 8 coroutines for the 9B model, with code released at github.com/dvmazur/async_llm.

Why it matters

LLMs currently process one thing at a time: a model generating a response cannot react to new input until it finishes. That blocks voice assistants, embodied agents, and monitoring systems, which operate in environments where the world keeps moving. AsyncLLM addresses this for any model already in production, with no retraining cost. The DevOps result is an explicit tradeoff, not a clean win: roughly half the forward passes at around 6 percentage points lower accuracy, which may suit latency-sensitive deployments where exhaustive compute is not an option. The throughput result (106 to 552 tokens/sec, 1 to 8 coroutines, single H200, 9B model) suggests the batching overhead is low enough to be useful in practice.