The world still contains vast amounts of unused data. But the cheap, clean and permissionless text that powered the first LLM boom is becoming polluted by AI output, contested by its owners and costly to replace. This week, AI companies were reportedly buying old books while Nvidia released a simulator that teaches robots through video, motion and synthetic consequences.

Get more from AI Weekly

More signal, less noise — pick your channels.

You're reading the weekly brief. Below are the other ways to follow the story — every channel free, easy to leave.

  • → Explore 16 deep dives
    Weekly topic-specific newsletters: Generative AI, Machine Learning, AI in Business, Robotics, Frontier Research, Geopolitics, Healthcare, and more.
    Browse all 16 deep dives →
  • → Breaking AI alerts
    Important developments that happen after your morning Espresso, without repeating what you already read. Usually no extra email; at most one afternoon update, plus a rare critical exception.
    Get breaking alerts →
  • → AI News Today (live)
    Live dashboard updated as the scanner finds news: scored stories from the last 48 hours, weekly entity movers, and quarterly trend lines across 113 AI companies, people, and topics.
    Open AI News Today →

In the Wild

Three links the expert network surfaced that we have not already sent you in Espresso.

  • Children may be AI’s first anti-hype demographic. Wired found kids describing generative AI as creepy, disgusting and uncool. Adoption is not only a capability problem; it is becoming a question of taste and identity. Read what the kids said
  • Wikipedia is designing an immune response to AI copy. Its new process lets editors remove suspected AI-generated contributions when defined conditions are met; anyone restoring them assumes responsibility for reviewing the content and its sources. Read the process
  • LLMs may pull human expression toward the same center. A cross-disciplinary review argues that widespread use can reinforce dominant styles and marginalize alternative voices. The risk is not only models learning from themselves; it is people beginning to sound more alike. Read the review

Quick Hits

The Open Frontier

Kimi K3’s weights are downloadable. Opus 5 shows a separate route to cheaper access.

The Compute Land Grab

The models get cheaper to use only after somebody commits the hardware, power and capital.

Who Pays the Bill

The grid and compliance costs are becoming explicit political choices.

Into Sensitive Systems

A share link can be technically public while users still experience it as private.

What happens when AI runs out of clean human text?

Last week’s old-books story looked like an odd procurement detail. Put it beside the next three links and it becomes a map of the post-crawl AI economy.

  • Old books become AI inventory. AI companies are reportedly buying printed books because they are guaranteed to predate AI-generated content. That is not nostalgia; clean human text now has procurement value.
  • xAI fights the ingredient label. California asks for general documentation of training datasets; xAI is challenging the rule. The corpus recipe is now competitive information.
  • Agents begin writing their own curriculum. Skill Self-Play has agents generate tasks and verify the results. Training material becomes something the system helps manufacture, not a fixed pile it eventually exhausts.
  • Nvidia turns movement into training data. Cosmos-H-Dreams learns from surgical video and robot kinematics, then simulates what happens next. The new “document” is an environment with consequences.

The point: the next data moat is not “more content.” It is control over reliable experience—licensed human archives, verifiable synthetic practice or proprietary physical-world environments. Publishers, simulator builders and robot operators become part of the model stack.

Key Takeaways

  • The model market is splitting in two: Kimi K3 expands what can be downloaded, while Opus 5 lowers the price of closed-model access.
  • Cheap intelligence rests on expensive infrastructure. The AMD and Nvidia commitments make the compute race visible; the ratepayer pledge asks who absorbs its external cost.
  • Trust still breaks at the default. Claude users created public share links, but search engines made “public” far more discoverable than many expected.
  • Upstream, the scarce asset is becoming reliable experience—not raw volume. That is where the next durable AI advantage may sit.

Worth Reading

Read the argument, not another recap: one analysis explains why Kimi K3 pressures closed-model economics; Anthropic draws the safety line it wants around powerful releases.

Worth Watching

Two fresh expert-shared videos, curated on AI TV.

Wait, What?

This week’s poll

Which source will matter most for the next jump in AI capability?

Last week, 131 of you voted:

After this week’s containment failures, where would you spend the next AI-security dollar?

  • Stronger sandboxes and access controls23%
  • Continuous AI-powered code scanning40%
  • More human review and incident response18%
  • Independent testing and disclosure rules19%

See full results →

Back Friday.

Alexis