huggingface.co web signal

FLaRe Hits 97% of CoT Accuracy on GSM8K at 3.9x Speedup

Generative AI Inference ai-business

TL;DR

  • FLaRe reaches 97% of explicit chain-of-thought accuracy on GSM8K while running 3.9x faster than token-by-token reasoning, per the paper.
  • The method compresses symbolic chains of thought into 8 latent slots of 512 dimensions, then generates them from Gaussian noise via flow matching.
  • Two-stage training: VAE and flow on 385K GSM8K-Aug examples, then self-training on verified rollouts across roughly 100K more questions on 16 H100 GPUs.

The paper reports that FLaRe, short for Flow-based Latent Reasoning, reaches 97% of explicit chain-of-thought accuracy on GSM8K while running 3.9x faster than reasoning out loud.

Latent reasoning, the authors write, "lets a large language model (LLM) think in a continuous space and verbalize only the answer." They argue a usable latent thought has to be five things at once — useful, diverse, explainable, refinable, and efficient — and build the system around a VAE that compresses symbolic chains of thought into 8 slots of 512 dimensions, paired with flow matching that draws those thoughts from Gaussian noise.

Training comes in two stages. First a VAE and a flow model fit on 385K symbolic CoT examples from GSM8K-Aug; then self-training on verified rollouts across roughly 100K additional questions, with answer loss backpropagated through the full rollout. Direct-read accuracy lands at 59.1% after stage two, against an explicit-CoT baseline of around 61% on Llama-3.2-1B. Training ran on 16 H100 GPUs with bf16 precision.

Against prior latent-reasoning baselines the paper puts FLaRe at a "3.9x speedup at 97% accuracy," versus CODI at 2.0x at 94% and PCCoT at 3.9x at 91%. It is one of several latent-generation papers on our radar this week, alongside EVA's VAE-prior swap. The paper was posted to Hugging Face on October 6.