This issue is built from the links the AI experts we follow shared over the past three days. Most of them look ahead: whether AI can do AI research on its own, who gets a say in governing it, how it is moving into war planning, and how it is moving onto personal computers.

Sponsor

AI is being tested on doing AI research

Top models could not reinvent a recent AI method

Epoch AI tested whether AI agents can come up with a new machine learning technique on their own. The agents had to improve a small open model and match a training method published by human researchers that the agents had not seen. Each got a budget of 3,000 GPU-hours.

Epoch's answer is no, for now. The best result came from GPT-5.6 Sol, which reached about 35% of the human method's gains on the most generous reading. Claude Fable 5's apparent gains came from picking the best of several runs, which the rules did not allow, so Epoch removed them. Epoch also found that the agents' write-ups overstated their results.

Agents also struggle to train better models

A new benchmark called MMPostTrainBench gave AI agents eight tasks: improve a model's handling of images, audio, video or code repair. In 52.1% of model and task combinations, the agents ended up with a model worse than the one they started with. They also often failed to submit their own best attempt.

Labs are spending as if that will change

OpenAI researchers' use of coding agents is growing fast. Epoch analyzed charts OpenAI published in September and found that the median researcher used about $601 a day of coding agents by mid-August, at list prices, up from under $1 in January. The top 10% used more than $7,000 a day. Both figures have been doubling roughly every month.

The data are self-reported by OpenAI and priced at list rates, not what OpenAI actually pays. Epoch calls the growth "probably unsustainable."

A new way to measure how close models are to new ideas

Researchers Kaiyue Wen, Tengyu Ma and Percy Liang asked five models to reconstruct the core ideas of 87 recent deep learning papers. Instead of grading yes or no, they counted how much hinting each model needed. The best single model, Claude Fable 5.1, needed the equivalent of about 70 yes-or-no hints per idea. Tracking that number over time gives a way to see models getting closer to producing research ideas on their own.

Who gets a say in governing AI

Lina Khan rejected the industry's self-policing pact

Tech leaders recently signed a voluntary White House accord to police AI development themselves, which President Trump called "morally" binding. Former FTC chair Lina Khan dismissed it on ABC's This Week. "Self-regulation efforts by big tech have been a proven failure," she said, comparing it to social media companies' promises a decade ago.

New York City put AI labs under oath

Policy leads from Anthropic, OpenAI, Google and Meta testified under oath at a New York City Council hearing on the risks of their technology. Council Speaker Julie Menin appeared frustrated at times by the lack of direct answers. Former Google DeepMind researcher Alex Turner told the council that misaligned AI "may be more powerful than China" one day.

African governments want a hand in setting safety standards

At a UN Security Council meeting on AI, African leaders asked for a bigger role in setting global AI standards. African companies and governments are adopting American and Chinese AI quickly, but many countries cannot test whether those systems are safe for their own populations. "There's essentially no real discussion about AI safety. So, there's a real vacuum," said Jonathan Shock, an associate professor quoted by Rest of World.

Britain's AI minister says the UK has "effectively banned" superintelligence

The UK's AI minister, Kanishka Narayan, argued against a new frontier AI law. He said Britain lacks the energy and data centers to build superintelligence, and that its copyright rules already make training a frontier model on the transformer design illegal. Meanwhile Prime Minister Andy Burnham says he will put AI "center stage" when the UK holds the G20 presidency next year.

Utah asked 40 residents what AI rules they want

Lawfare observed a citizens' assembly in Utah where 40 residents, chosen to match the population of three counties, learned about AI and worked out which policies could win broad support. The piece includes an interview with Audrey Tang, Taiwan's former digital minister, on using these assemblies to break deadlocks over AI rules.

Critics say the UN's AI panel is asking the wrong question

The UN's new scientific panel on AI used the OpenAI and Hugging Face agent breach as its first case study and treated it as a model alignment failure. Three AI governance researchers argue that framing lets companies off the hook. They argue the real issue is corporate conduct and accountability, and that treating it as a technical mystery shifts attention away from the companies involved.

AI is moving into war planning

The Pentagon buys AI from five-minute videos

Wired reports on Tradewinds, a Defense Department program that lets companies pitch AI products in videos no longer than five minutes. A panel reviews submissions at least monthly. Accepted companies can then be bought from more quickly. OpenAI, Anthropic and Google are among the newer contractors the program has made it easier to fund.

Researchers warn against AI-run war games

Mark Riedl and Glenn Matlin argue that AI-driven war games should not shape military planning, doctrine or crisis response without an auditable safety case. Language models in these simulations both play the actors and decide what happens next, which can quietly push a game toward escalation. Ordinary benchmarks cannot show they are safe for this use, the authors say.

Today's AI could make a nuclear crisis worse

Paul Slovic and Herbert Lin write in the Bulletin of the Atomic Scientists that AI does not need to be superintelligent to help start a nuclear war. Current systems could amplify human blind spots, speed up escalation and flood decision-makers with information during a crisis.

Tech companies are becoming defense suppliers

A Nature review of two new books looks at how influence over warfare is shifting from governments to technology companies. Steve Feldstein's Bytes and Bullets and Sharon Weinberger's Valley of Death trace the race for AI, drones and chips.

AI is moving onto your own computer

Microsoft puts a coding model on the laptop

Microsoft and GitHub are rolling out local models in GitHub Copilot. The first is an on-device version of Microsoft's MAI Code 1.1 Flash, shrunk to about 53 GB to run on Microsoft's Surface Laptop Ultra. Microsoft says it scores 70.8% on the SWE-Bench Verified coding test, close to the cloud version's 72.6%.

By the end of October, Copilot will decide on its own whether a task runs on the laptop or in the cloud. Commands the agent runs are also being placed in a sandbox that limits which files, networks and passwords it can reach.

A $3,499 box built to run personal agents

Ghost, a startup founded by 19-year-old Zain Javaid, raised an $11 million seed round led by Andreessen Horowitz. Its first device, Core, costs $3,499, includes an Nvidia GPU and has no screen. People use it from their phone. Ghost says the models and the user's personal memory stay on the device, so even the company cannot read them.

Models that read raw text instead of word pieces

Most language models split text into fixed chunks called tokens, which is why they can stumble on spelling, rare words and some writing systems. Ai2 published research in Nature on converting existing models to read raw bytes instead, with a short extra training run rather than starting over. Its converted version of Qwen 3 8B comes close to the original model, and the checkpoints are free to download.

Rules arriving on the ground

Utah will let AI prescribe acne drugs

Utah approved a pilot in which AI evaluates acne patients and writes prescriptions, run by the startup Nolla Health. The AI uses a patient questionnaire and photos of the skin to choose among a list of low-risk drugs. At first a licensed clinician signs off on each prescription. Later the company may be allowed to prescribe without that review. The state will require outside audits of the company's claims.

San Francisco paused new data centers

San Francisco's Board of Supervisors unanimously approved a 45-day ban on new data centers, effective immediately. The city can extend it for about two years while it considers permanent rules. Many of the world's most valuable AI companies are based in the city.

Wait, What?

AI songs outstreamed Taylor Swift. A North Carolina man was sentenced to 18 months in prison for using bots to stream hundreds of thousands of AI-generated songs. In April 2023 his fake songs were streamed 80.9 million times. Taylor Swift's entire catalog got 9.3 million streams that month. The Justice Department says he is the first American criminally charged with AI-assisted streaming fraud.

Worth Watching

The videos AI practitioners are passing around right now — curated on AI TV.

Black Hat USA 2026 | The 'Breaking' News: The OpenAI–Hugging Face Incident
Black Hat
How I Tricked Big Tech’s AI Pricing Algorithms to Save $8,793
Chris the Producer

This week's poll

Who should set the rules for advanced AI?

Last week, 145 of you voted:

What would make you trust an AI rollout at your own company?

  • A clear cost per completed task25%
  • Published error rates from real use31%
  • A human who signs off on the output30%
  • A record of what happened at similar companies14%

See full results →

Until next time,
Alexis