Chelsea Finn

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
1
past 30d
Sources
1
distinct domains
Discussões
0
past 30d
Latest signal
29d ago
View every signal from Chelsea Finn →

Articles & links

LLM RL optimizes for sequential reasoning We also optimize over the reasoning strategy, incl parallel trains of thought, aggregation of parallel traces, & sequential reasoning This allows the model to better explore & allocate compute at test time https://t.co/na0GAxbY…

SPIRAL: Learning to Search and Aggregate arxiv.org
AI Weekly's analysis
  • SPIRAL co-trains three reasoning primitives in one RL framework: sequential chain-of-thought, parallel sampling of traces, and learned aggregation of those traces.
  • The paper reports outperforming GRPO by up to 11× scaling efficiency and 15% higher performance when all three compute primitives are scaled.
  • Training uses set reinforcement learning to make parallel traces collectively useful, plus standard RL to train the aggregation step itself.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 72d ago

Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch. We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining. Paper: https://t.co/dlj5RXVFED https://t.co/hJHpFes03y

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? arxiv.org
AI Weekly's analysis
  • Naive Q-function pretraining on offline data often provides little benefit over random initialization when fine-tuning a pretrained policy online.
  • The mismatch: pretraining targets the pretrained policy's Q-function, not the Q-function that online fine-tuning actually converges to.
  • The proposed IPE method trains multiple diverse policies and pools their rollouts, yielding a 1.26x average improvement on continuous control benchmarks.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 29d ago

Are you Chelsea Finn? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/chelsea-finn)