arxiv.org web signal

Technion team: LLM leaderboards flip when prompts rephrased

Safety Hallucinations ai-business

TL;DR

  • A Technion team measures LLM generalization as stability under paraphrase across four behavioral axes, not aggregate accuracy on a benchmark.
  • The paper finds no model generalizes uniformly across eleven tested LLMs, and cross-dataset variation can reverse rankings on six benchmarks.
  • G4-E4B records the open-source pool's highest generation instability (Δ-BLEU 0.837); Q3.5-9B records the lowest, but the highest confidence instability.

Model rankings flip when the dataset changes, according to a NeurIPS 2026 workshop paper from a Technion team that re-defines generalization as behavioral stability under paraphrase rather than accuracy on a benchmark. The group, composed of Nagham Omar, Mahmoud Jabarin, Maya Rozensztein, Rom Himelstein, Avi Mendelson and Amit LeVi, tested eight open models and three closed ones (GPT-5.4, Gemini-2.5-Flash, Claude-Sonnet-4.6) across six datasets, introducing a framework called the Stability-Aware Generalization Objective, or SAGO. The arxiv paper, accepted at the TAE "Can We Trust AI Evaluation?" workshop, measures how a model's behavior shifts when the same input is rephrased, across four axes: generation consistency, internal activations, confidence and response mirroring.

The paper's core claim, verbatim: "no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings." The pairing the authors highlight makes the point concrete. G4-E4B records the highest generation instability in the open-source pool (Δ-BLEU 0.837), while Q3.5-9B records the lowest (Δ-BLEU 0.520) but the highest confidence instability (Δ-Prob 0.088). Which axis you pick decides which model looks worse.

Scaling up does not fix it. The paper reports that "L-8B, G-7B, and Q-7B are equally or more unstable than their smaller counterparts on most axes." Nor does paying for a frontier API: "The open- versus closed-source differences further indicate that stronger models do not eliminate instability, but may shift it to different behavioral dimensions."

Within a single model, the numbers also move from dataset to dataset. For Llama-3.1-8B, "Δ-Cos spans a 3.7× range and Δ-MR more than doubles across the six datasets." The authors' summary: "A model appearing stable under one evaluation may be substantially less stable in practice."