arxiv.org web signal

Multilingual prompting unlocks LLM facts across 17 languages

TL;DR

  • The paper evaluates cross-lingual prompting on multilingual factual benchmarks covering 17 typologically diverse languages.
  • Cross-lingual exploration shows a more efficient compute Pareto frontier than native-language scaling.
  • Gains in cross-lingual consistency exceed what accuracy improvements alone would explain.

Prompting a large language model in several languages at once surfaces facts the model would otherwise leave hidden, according to a preprint by Elisha Diskind, Itamar Trainin, Uri Shaham, Leshem Choshen, Idan Szpektor and Omri Abend.

The team evaluated cross-lingual prompting strategies on multilingual factual benchmarks covering 17 typologically diverse languages, and identifies "four inherent dimensions of cross-lingual exploration that directly govern parametric knowledge retrieval." Two of the researchers we follow had already circulated the link when we picked it up.

The headline claim is about efficiency, not just accuracy. The approach, the authors write, represents "a more efficient compute Pareto frontier than native-language scaling" — inference-time prompt diversity buying more factual recall per unit of compute than pouring the same budget into monolingual scaling. Improvements in cross-lingual consistency, they add, exceed "what can be explained by accuracy gains alone."

The published abstract names neither the specific benchmarks used nor the four dimensions themselves, and reports no per-language numbers.

Shared on Bluesky by 2 AI experts