FLaRe Hits 97% of CoT Accuracy on GSM8K at 3.9x Speedup
TL;DR
- FLaRe reaches 97% of explicit chain-of-thought accuracy on GSM8K while running 3.9x faster than token-by-token reasoning, per the paper.
- The method compresses symbolic chains of thought into 8 latent slots of 512 dimensions, then generates them from Gaussian noise via flow matching.
- Two-stage training: VAE and flow on 385K GSM8K-Aug examples, then self-training on verified rollouts across roughly 100K more questions on 16 H100 GPUs.
The paper reports that FLaRe, short for Flow-based Latent Reasoning, reaches 97% of explicit chain-of-thought accuracy on GSM8K while running 3.9x faster than reasoning out loud.
Latent reasoning, the authors write, "lets a large language model (LLM) think in a continuous space and verbalize only the answer." They argue a usable latent thought has to be five things at once — useful, diverse, explainable, refinable, and efficient — and build the system around a VAE that compresses symbolic chains of thought into 8 slots of 512 dimensions, paired with flow matching that draws those thoughts from Gaussian noise.
Training comes in two stages. First a VAE and a flow model fit on 385K symbolic CoT examples from GSM8K-Aug; then self-training on verified rollouts across roughly 100K additional questions, with answer loss backpropagated through the full rollout. Direct-read accuracy lands at 59.1% after stage two, against an explicit-CoT baseline of around 61% on Llama-3.2-1B. Training ran on 16 H100 GPUs with bf16 precision.
Against prior latent-reasoning baselines the paper puts FLaRe at a "3.9x speedup at 97% accuracy," versus CODI at 2.0x at 94% and PCCoT at 3.9x at 91%. It is one of several latent-generation papers on our radar this week, alongside EVA's VAE-prior swap. The paper was posted to Hugging Face on October 6.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper 'What Matters for Latent Reasoning' Introduces FLaRe, Hits 97% CoT Accuracy at 4x Speedup