The fun number needs a footnote: Claude Opus 5 scored about 30% alone, then reached 100.00 RHAE inside Nvidia's AVO harness on all 183 public levels. Persistent memory, a supervisor, and tool use did the lifting. AVO used 12% fewer environment actions than VISTA, though Nvidia says that is not a controlled comparison. The private set remains the test to watch.
AI news for Saturday, August 22, 2026
The Daily AI Espresso — the links the most-followed people in AI actually shared, curated every morning. Edited by Alexis · Live updates →
Three consequential moves made Absolute Top Alerts: Nvidia's public-benchmark leap, Amazon's AI-linked device price hikes, and Anthropic's reported $2 trillion IPO target.
The base Echo Dot jumped from $49.99 to $79.99; a 16GB Kindle from $109.99 to $149.99. Amazon confirmed the increases and cited memory and storage costs. The wider squeeze comes from AI compute demand—and it is now landing in ordinary shopping carts.
These are banker-sourced possibilities, not an Anthropic price tag. Still, a raise larger than most technology companies at a valuation near the top of public markets would turn frontier-model economics into a mass-market bet almost overnight.
A lighter Saturday tour through models, science, policy, tools, and the odd corners where AI culture is forming.
Brazil is putting 1.3 billion reais into a Huawei-iFlytek project in Rio and 1 billion into a separate supercomputer tender expected—but not guaranteed—to favor Nvidia. The explicit strategy: control national data without depending on one company, technology, or country.
Sol now costs $4 per million input tokens and $20 per million output tokens: 20% and 33.3% cuts respectively. The promotion lasts at least through November 21. If you run long coding or research jobs, reprice them—but do not build a 2027 budget around a temporary sticker.
Harvey Tenet starts from Kimi K3 and is post-trained with Fireworks for long-horizon legal work. Harvey says it nearly doubled completed LAB hold-out tasks and reached the top of LAB Contracts while keeping cost stable. It is a research preview, and the benchmark claims are Harvey's—but the open-weight, no-customer-data route is notable.
The analysis estimates AI-assisted writing signals in 77% of English PubMed Central papers across 2025 and almost nine in ten that December. Introductions and discussions showed more signs than methods and results. Huge caveat: this is a word-pattern estimate from a preprint, not proof that 90% of the science was generated or wrong.
OpenAI is asking its home state to move safeguards upstream: monitor frontier models while they are being trained, not only when they approach release. California has not adopted the proposal. The reversal is the signal—recent agent incidents are turning voluntary lab practice into requests for enforceable rules.
DeepMind traces the road from Atari and AlphaGo to SIMA 2, then points at its current EVE Universe partnership. Persistent player economies add diplomacy, conflict, memory, and surprise—the messy ingredients missing from neat benchmarks. The weekend question: can a useful general agent learn to live in New Eden first?
What is buzzing across X, Reddit, and Hacker News—labeled before the hype outruns the evidence.
OpenRouter's free Ox Alpha preview has a one-million-token context window, image and video inputs, and no named developer. Weekend speculation calls it Gemini 3.5 Pro. One detailed community fingerprinting attempt reports a GLM-matching tokenizer, Z.ai-style errors, and nearly identical temperature-zero outputs—so the strongest clues currently lean GLM. Nobody has claimed it. A fun mystery, not a procurement plan.
NoBuzz adds a /debuzz skill to Claude Code, then asks Gemini to rewrite the last answer for a colleague, manager, or director. Yes: one model cleans up another model's TED-talk voice. The MIT-licensed joke also lands a useful product lesson—tone and audience fit are becoming their own tool layer.
That’s the weekend shot. — Alexis · AI Weekly