Out of curiosity I had Fable write a response to the article about the weaknesses of LLMs: www.nytimes.com/2026/06/30/o... It's uh, pretty good. claude.ai/code/artifac...
Eugene Vinitsky
Reinforcement learning and autonomy researcher
Reinforcement learning and autonomy researcher with public evidence across Agents & robotics, AI research.
- AI signals
- 9 past 30d
- Sources
- 5 distinct domains
- Discusiones
- 90 past 30d
- Latest signal
- 8d ago
Articles & links
Fantastic paper demonstrating how LLM editing is gradually distorting our writing and the way we think to write: arxiv.org/abs/2603.18161
- Heavy LLM users produced nearly 70% more essays that stayed neutral on the topic question, and reported the writing felt less creative and not their own.
- Even when asked only to fix grammar in 2021 human-written essays, LLMs significantly altered the semantic meaning of the text.
- About 21% of peer reviews at a recent top AI conference were AI-generated, and those reviews scored papers a full point higher on average.
LLMs will likely be the most powerful propaganda tool that has ever existed arxiv.org/abs/2606.16475
- Across 18,978 conversations with 6,923 people, AI systems reliably out-persuaded expert humans, including world championship debaters and professional canvassers.
- Experts chose their topics, researched in advance, went through hours of structured practice, and were paid £1,000 cash bonuses, and still lost to AI.
- In a live fundraising test for Save the Children, AI was nearly 3x more effective than professional canvassers at raising real donations.
This is a really excellent benchmark, pointing to some serious missing capabilities in coding agents: arxiv.org/abs/2603.247.... They don't write code with an eye towards maintenance!
- SlopCodeBench tested 11 models on 20 problems across 93 checkpoints; no agent solved any problem end-to-end and the top checkpoint solve rate was 17.2%.
- Structural erosion rose in 80% of trajectories and verbosity in 89.8%; agent code was 2.2x more verbose than 48 open-source Python repositories.
- Human code stayed flat over time while agent code deteriorated each iteration, and prompt interventions lifted initial quality but did not halt degradation.
New Paper: arxiv.org/abs/2606.19370 Self-play yields capabilities but requires frustrating cost-function tuning. Surprisingly, just 30 minutes of demonstration data produces much more human-like driving policies! Led by @daphne-cornelisse.bsky.social Website: spiced-self-play.com
- A new method trains driving AI using only 30 minutes of human demonstrations, 2,500 times fewer than comparable imitation learning approaches.
- Human demonstrations serve as a regularization signal on top of a basic goal-reaching reward, not as the primary training objective.
- Resulting policies coordinate with held-out human trajectories and finish training in 15 hours on a single consumer-grade GPU.
These are very odd failures! arxiv.org/abs/2211.00241
- An adversarial policy achieves a greater than 97% win rate against KataGo running at superhuman settings.
- The attack transfers zero-shot to other superhuman Go AIs and works even against KataGo variants adversarially trained to defend.
- Human experts can execute the winning strategy manually, defeating superhuman Go AIs without any computer assistance.
Does physical modeling matter or should we just do data-driven stuff? I feel like this paper characterizes things perfectly: arxiv.org/abs/2410.23179. Physics-based modeling is always helpful but the relative gain closes at scale
- In rigid-body interaction experiments with transformers, equivariant models beat non-equivariant ones at every tested compute budget.
- Non-equivariant networks trained with data augmentation can close the data-efficiency gap, but only given sufficient training epochs.
- Optimal compute allocation between model size and training steps differs significantly between equivariant and non-equivariant architectures.
Would be interesting to see this offered as a service at conferences and journals...https://arxiv.org/abs/2608.13331
- Faraday, a 27B-parameter AI agent, outperformed Claude Opus 4.8 and GPT-5.5 on held-out paper replication tasks.
- The team built Replica, a scalable task space for replication, plus an auto-generated rubric-based judge said to track human assessment.
- Authors frame Faraday as a stepping stone toward AI agents doing long-horizon scientific work without complex harnesses.
Recent commentary
A lot of my fears about LLM usage come from using Stackoverflow when I was a wee programmer. It was easy to fall into copy pasting behavior that long term made you worse
Repulsed by the idea that we should sit idly by while frontier labs determine the trajectory of the AI transition. They don’t know how to make systems “safe” or “aligned“ nor do any such notions make sense without democratic input. There’s so much research to do here.
My position on AI consciousness is that you really should try to eat less meat if you can
Taking more seriously the claim that recent ML models do not "reason", it still is quite odd the particular ways that superhuman game-playing models fail, suggesting that what they're doing is something quite different from human reasoning about play
Very obviously the outcome of people mobbing those who publicly admit they use AI is everyone continuing to use it but hiding it
"LLMs mean you don't have to know things, you can just do things" assumes a clean separation that probably doesn't exist. You know what you want to do in part because of what you know / understand
I'm not worried about people mistreating AI if it ever seems conscious. Humans have an incredible history of taking good care of conscious beings. After all, 99% of humans are vegetarian.
Proposed norm: you're allowed to respond to an unedited LLM document with an unedited LLM document that is twice as long
(1) Artists are justified in being angry about AI and we really shouldn’t use AI art where otherwise we would have hired an artist. Art isn’t really about productivity growth and automating it helps no one. (2) Programming is different; we are bottlenecked on software and need more of it urgently
I have already met some number of people who are basically a front-end to an LLM
In Eugene Vinitsky's orbit
Center = Eugene Vinitsky. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Eugene Vinitsky? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/eugene-vinitsky)