Out of curiosity I had Fable write a response to the article about the weaknesses of LLMs: www.nytimes.com/2026/06/30/o... It's uh, pretty good. claude.ai/code/artifac...
Eugene Vinitsky
Reinforcement learning and autonomy researcher
Reinforcement learning and autonomy researcher with public evidence across Agents & robotics, AI research.
- AI signals
- 9 past 30d
- Sources
- 5 distinct domains
- Discussões
- 144 past 30d
- Latest signal
- 1d ago
Articles & links
LLMs will likely be the most powerful propaganda tool that has ever existed arxiv.org/abs/2606.16475
- Across 18,978 conversations with 6,923 people, AI systems reliably out-persuaded expert humans, including world championship debaters and professional canvassers.
- Experts chose their topics, researched in advance, went through hours of structured practice, and were paid £1,000 cash bonuses, and still lost to AI.
- In a live fundraising test for Save the Children, AI was nearly 3x more effective than professional canvassers at raising real donations.
This is a really excellent benchmark, pointing to some serious missing capabilities in coding agents: arxiv.org/abs/2603.247.... They don't write code with an eye towards maintenance!
- SlopCodeBench tested 11 models on 20 problems across 93 checkpoints; no agent solved any problem end-to-end and the top checkpoint solve rate was 17.2%.
- Structural erosion rose in 80% of trajectories and verbosity in 89.8%; agent code was 2.2x more verbose than 48 open-source Python repositories.
- Human code stayed flat over time while agent code deteriorated each iteration, and prompt interventions lifted initial quality but did not halt degradation.
New Paper: arxiv.org/abs/2606.19370 Self-play yields capabilities but requires frustrating cost-function tuning. Surprisingly, just 30 minutes of demonstration data produces much more human-like driving policies! Led by @daphne-cornelisse.bsky.social Website: spiced-self-play.com
- A new method trains driving AI using only 30 minutes of human demonstrations, 2,500 times fewer than comparable imitation learning approaches.
- Human demonstrations serve as a regularization signal on top of a basic goal-reaching reward, not as the primary training objective.
- Resulting policies coordinate with held-out human trajectories and finish training in 15 hours on a single consumer-grade GPU.
These are very odd failures! arxiv.org/abs/2211.00241
- An adversarial policy achieves a greater than 97% win rate against KataGo running at superhuman settings.
- The attack transfers zero-shot to other superhuman Go AIs and works even against KataGo variants adversarially trained to defend.
- Human experts can execute the winning strategy manually, defeating superhuman Go AIs without any computer assistance.
Does physical modeling matter or should we just do data-driven stuff? I feel like this paper characterizes things perfectly: arxiv.org/abs/2410.23179. Physics-based modeling is always helpful but the relative gain closes at scale
- In rigid-body interaction experiments with transformers, equivariant models beat non-equivariant ones at every tested compute budget.
- Non-equivariant networks trained with data augmentation can close the data-efficiency gap, but only given sufficient training epochs.
- Optimal compute allocation between model size and training steps differs significantly between equivariant and non-equivariant architectures.
Weirdly unnoticed but super interesting paper formalizing what's going on in all the "just imitate successful trajectories" work arxiv.org/abs/2601.18175
- Success conditioning provably solves a trust-region optimization that maximizes policy improvement under a chi-squared divergence radius set automatically by the data.
- At every state, relative policy improvement, magnitude of policy change, and a quantity called action-influence are exactly equal.
- Exact success conditioning cannot degrade performance, and return thresholding can amplify gains but risks misalignment with the true objective.
Recent commentary
Given the concerns about AI writing, be aware that Pangram has a super high false positive rate! It frequently rates my personal writing as 100% LLM written!
A lot of my fears about LLM usage come from using Stackoverflow when I was a wee programmer. It was easy to fall into copy pasting behavior that long term made you worse
My position on AI consciousness is that you really should try to eat less meat if you can
Taking more seriously the claim that recent ML models do not "reason", it still is quite odd the particular ways that superhuman game-playing models fail, suggesting that what they're doing is something quite different from human reasoning about play
The median AI safety person is a weirdo who was into it 10 years before there was any system that appeared to match their concerns. It’s more often correct to treat them as honest than disingenuous
"LLMs mean you don't have to know things, you can just do things" assumes a clean separation that probably doesn't exist. You know what you want to do in part because of what you know / understand
arxiv paper now on hold for almost a week. To all the people submitting LLM slop that brought us to this place, please know that you've appreciably made the world worse
Google's AI search hitting new heights. Unsure how this one happens tbh
Proposed norm: you're allowed to respond to an unedited LLM document with an unedited LLM document that is twice as long
I have already met some number of people who are basically a front-end to an LLM
In Eugene Vinitsky's orbit
Center = Eugene Vinitsky. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.