paper web signal

NYU: LLM chain-of-thought 44.8–75.9% faithful across tasks

TL;DR

  • CoT traces matched internal computation only 44.8%–75.9% of the time across three tasks and three mid-sized open models, NYU researchers report.
  • The team introduces CoT-Interpretability Alignment (CIA), using linear probes to compare a model's spoken reasoning with the strategy encoded in its representations.
  • Post-training that rewards both task accuracy and parametric faithfulness delivered an average relative gain of 25.5% while preserving task performance.

Three mid-sized LLMs produce chain-of-thought traces that align with their internal computation only 44.8% to 75.9% of the time, according to a new paper from Yihuai Hong, Shauli Ravfogel, Chen Zhao and Eunsol Choi at New York University and NYU Shanghai.

The researchers test Llama3.1-8B-Instruct, Gemma2-9B-it and Qwen3-8B on three tasks: two-hop factual reasoning, hint interventions and integer multiplication. Their metric, CoT-Interpretability Alignment (CIA), uses linear probes to read what strategy the model's internal representations encode, then checks whether that strategy shows up in the model's spoken reasoning. The abstract warns that "models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers."

Post-training with task accuracy and parametric faithfulness as rewards yields an "average relative gain of 25.5%" with the best method per setup, while preserving or improving task accuracy. The authors conclude that "CoT parametric faithfulness is both measurable and improvable, offering a pathway toward more trustworthy explicit reasoning in LLMs."

The abstract publishes no per-task or per-model breakdown of the 44.8–75.9% range; code and data are posted on GitHub.

Shared on Bluesky by 1 AI expert