Found first: a primary source the press has not covered yet.
A new paper on arXiv finds that a prominent non-invasive brain-to-text result achieves nearly identical scores on signals carrying zero neural information as on real brain recordings: 22.0% balanced accuracy versus 22.3%. Dulhan Jayalath and Oiwi Parker Jones trace the near-equivalence to a timing artifact in how the prior method encodes input windows, and show that fixing it produces genuinely better decoding. The paper is arXiv:2609.40359.
What the source says
Jayalath and Jones re-examine the method of d'Ascoli et al. (2025), which jointly encodes all overlapping windows in a sentence. Because consecutive windows overlap around word onsets, their joint encoding implicitly reveals word duration, giving the model a timing shortcut it can exploit in place of any neural signal. Their proposed fix, SimpleB2T, processes each window independently, removing that shortcut. On perceived-speech benchmarks, SimpleB2T reaches 36.6% word error rate using five observations per word, a level the authors describe as approaching past invasive speech decoding performance. Two existing strategies, aggregating predictions across multiple neural responses to the same word and applying pretrained language model priors, become substantially more effective once the shortcut is gone.
Why it matters
The gap between 22.0% on synthetic signals and 22.3% on real brain recordings is narrow enough to be indistinguishable from noise, which means the prior method's reported gains were driven almost entirely by word-timing patterns available without any neural measurement. That raises questions about whether the same artifact exists in other non-invasive decoding pipelines that use overlapping temporal windows, since the encoding pattern is common. SimpleB2T achieves better performance than the artifact-affected baseline while genuinely depending on brain data, which puts comparisons to invasive methods back on more solid ground.