Sci-VBench: video generators look realistic, fail science reasoning
Sci-VBench evaluates 16 frontier proprietary and open-source video models on 1,253 expert-annotated prompts spanning 60 subjects across Natural Science, Healthcare, Humanities and Engineering. Under a rubric-based protocol validated with non-expert humans and MLLM-as-judge, perceptual-quality scores cluster tightly across systems but Prompt Grounding and Scientific-Causal Correctness diverge sharply, with a pronounced proprietary/open-source gap. The authors from Zhejiang, UCAS, Tongji and Yale conclude visual realism has not yet translated into reliable causal or scientific modeling.
Evo-Bench measures if LLMs can improve their own agent harnesses
Evo-Bench is the first benchmark to isolate a model's ability to autonomously optimize its own agent harness across Search, Office and General domains, using auxiliary-task evolution to pick harness-sensitive tasks and stratified splits to prevent overfitting. Across nine frontier and open-weight models, top systems earn absolute gains up to 16.6 points, nearly matching human-engineered harnesses on General and Search tasks but stalling on Office workflows. The paper also flags early-saturation temporal anomalies during self-evolution.
Mind Lab's Macaron-V1 wires LoRA specialists onto a frozen 744B GLM-5.2
Macaron-V1 freezes a 744B GLM-5.2 base and composes four specialist LoRA adapters (chat, agent, coding, GenUI) that are selected per turn, with a smaller Macaron-V1-Tall variant built on a 50B Qwen3.6 model for local deployment. The paper pairs the mixture-of-LoRA architecture with a recursive model-harness co-design loop and the MindForge agentic RL framework. The authors flag that compounding gains from continual learning across their pipeline remain an open question.
Motif 3 ships 314B MoE with grouped-differential latent attention
The Motif 3 technical report describes a 314B-parameter sparse MoE that activates 13.2B parameters per token across 384 routed experts, with a novel Grouped Differential Latent Attention block and expert-specific polynomial activations. The model trains on ~12.5T tokens in MXFP8 with a 256K context window and post-trains via multi-teacher on-policy distillation. Reported base-model scores include 86.2% MMLU, 93.9% GSM8K and 73.7% HumanEval pass@1.
SWE-Bench ProMax debuts multilingual refactoring benchmark, top model 41.2%
SWE-Bench ProMax introduces 170 expert-curated, multilingual refactoring tasks across Python, Java, TypeScript, Go, C, C++ and Rust, averaging 11.4 modified files and 261.6 lines per instance. The authors rewrote every issue description and manually reviewed test suites to fix the overly-narrow and overly-broad tests that a recent audit found in 60% of unsolved SWE-Bench Verified instances. The best frontier model reaches only 41.2% resolve rate across two agent scaffolds.
theblock.co
31m ago
ALERTA 32
Anthropic locks 20-year, 191 MW compute deal with Riot for $9.1B
Riot Platforms disclosed a 20-year data-center lease at its Rockdale, Texas campus with a tenant that Bloomberg identifies as Anthropic. The 191 MW deal runs through June 2048 for ~$9.1B in base revenue, with two five-year extension options that could take total value to $16.1B. Phased delivery brings 96 MW online by December 2027 and the full 191 MW by June 2028; Morgan Stanley is providing $573M of interim financing, and RIOT shares jumped 25% after-hours.
enterpriseam.com
2h ago
20
Moody's flags bank AI dependence on OpenAI, Anthropic as systemic risk
Moody's published an Aug 10 warning that banks' rapid adoption of proprietary AI is creating a 'systemic dependency' on a concentrated group of loss-making tech suppliers, naming OpenAI and Anthropic specifically. The agency says both face investor pressure to reach profitability, giving them potential pricing leverage over dependent financial institutions, and adds that supervisors are likely to sharpen focus on operational resilience and vendor concentration. More than three-quarters of UK financial services firms already use AI, highest among insurers and international banks.
x.com
2h ago
23
Anthropic freezes Sonnet 5 at intro pricing, cancels Sept 1 hike
Anthropic posted on X that Claude Sonnet 5's June introductory pricing of $2 per million input tokens and $10 per million output tokens is now permanent, canceling the planned Sept 1 move to $3/$15. The freeze cuts token costs roughly 33% for input and 33% for output vs. the previously scheduled rates and applies to a model that ships with a new tokenizer that yielded ~30% more tokens than Sonnet 4.6 for the same text.
cryptobriefing.com
3h ago
23
Unitree Shanghai IPO draws 2,700x retail oversubscription
Unitree Robotics' Shanghai STAR Market IPO drew over 2,700x oversubscription from retail investors after subscriptions opened August 10, with Bloomberg later reporting 5,526x for retail buyers. The listing values China's first pure-play humanoid robot maker at approximately $9 billion, raising 6.1 billion yuan ($904 million) by selling 40.45 million new shares — 10% of enlarged capital — at 150.8 yuan each. DeepSeek is among strategic investors; CEO Wang Xingxing founded Unitree in 2016 after leaving DJI.
OpenAI's ethics chief Bakalar leaves after under a year
OpenAI's head of ethics Chloé Bakalar has left the company less than a year after joining, the Financial Times reported August 10. Her exit continues a wave of high-profile departures from OpenAI's safety and ethics leadership, including safety systems lead Johannes Heidecke in July and former safety systems head Lilian Weng. Departing researchers have publicly attributed exits to a perceived deprioritization of safety amid rapid commercialization and frequent model releases.
sbs.com.au
3h ago
25
South Australia unveils Australia's first AI royal commission
South Australian Premier Peter Malinauskas announced on August 10 a national-first royal commission into AI, warning unchecked development poses risks to society and employment. The inquiry begins in October and reports by July 1, 2027, examining AI's impact on work, health, education and data-center energy and water use, with terms of reference and three commissioners to be named within 4-6 weeks. The announcement followed Malinauskas's US trip meeting OpenAI, Anthropic and Apple.
Anthropic to watermark Claude output under EU AI Act
Anthropic says it has signed the EU AI Act Article 50(2) Code of Practice and is rolling out machine-readable marks on Claude outputs. All Claude models launched on or after 2026-08-02 support marking at launch, using imperceptible text watermarks plus C2PA metadata embedded in .svg/.png/.jpg files. Anthropic warns detection is probabilistic — marks can be lost through editing or format conversion, and their absence does not prove content is non-AI.
OpenAI backs Abbott's Texas data-center rules
OpenAI posted a letter to Governor Greg Abbott committing to comply with Texas's new data-center standards — pay for its own electric infrastructure, reuse water where possible, avoid disrupting residential neighborhoods, and forgo taxpayer-funded incentives. The letter lands days after Abbott froze new data-center power connections pending a PUC and ERCOT audit, and reframes OpenAI's Texas Stargate buildout as a good-faith compliance story rather than a grid-cost fight.
cactuscompute.com
7h ago
22
Cactus ships 14MB agentic LLM that runs on a Pi
Cactus Compute released Needle 2, a 45M-parameter, 14MB Apache-2.0 agentic model built for tool calling and device control on sub-$200 hardware. The company claims 500 tok/s decode on a Raspberry Pi 5 in about 28MB of RAM and 400-1,500 tok/s on a Meta Quest 3S, using 2-bit training-integrated quantization and a Walsh-Hadamard 'Simple Attention Network' instead of dense projections.
Applied Compute in talks at $3B, doubles in months
The Information reports Applied Compute — the ex-OpenAI enterprise-agent startup that raised $80M at a $1.3B valuation in April — is in talks to raise 'hundreds of millions' more led by Elad Gil at a roughly $3B valuation. The round would more than double the company's price in about four months as demand for custom fine-tuned models climbs.
OpenAI runs $7B employee tender at $852B, no bump
Bloomberg reports OpenAI ran a tender offer to buy back roughly $7B in shares from current and former employees, pricing the deal at the same $852B valuation set by its March 2026 primary round. The buyback keeps the mark unchanged rather than pushing it higher, an unusual signal from the most valuable private AI company as it prepares for a possible IPO.
blog.sshh.io
8h ago
19
Blog probes Claude and GPT knowledge cutoffs and training runs
Independent researcher Shrivu Shankar published a probing methodology on August 10 that infers training-run identity for frontier models from historical-fact quizzes, self-reported dates and self-identification. Findings: Anthropic's Opus 4.7+ models share a late-December-2025 cutoff (likely one training run), OpenAI's GPT-5.6 family clusters around a late-February-2026 checkpoint, and Opus 5 has an oddly older knowledge state (~January 2026) despite a published May 2026 cutoff. He also notes Anthropic models occasionally self-identify as GPT-4, suggesting training on prior-generation model outputs.
kuber.studio
8h ago
16
Essay argues humanising LLM outputs discards fidelity
Kuber Mehta's August 10 post argues that instructing LLMs to produce 'humanised' outputs — ADHD-friendly formatting, simplified language — pushes lossy compression into the middle of an agent pipeline rather than at the human boundary. The result, he argues, is silent information loss and obscured agent failures. Recommendation: keep highest-fidelity representations as long as possible and only transform at consumption time, the way databases and compilers do.
Ante ships single-binary offline Rust coding agent
Antigma Labs open-sourced Ante, a ~15MB Rust binary that ships as a self-contained terminal coding agent with embedded grep, git and llama.cpp for local GGUF inference — no external dependencies. The project reports 82.7% on Terminal-Bench 2.1 (89 tasks, 5 trials each) and claims roughly 7x less peak memory, 9x less average CPU and 5x less disk I/O than Claude Code. Source is Apache 2.0; the prebuilt binary is under a separate alpha 'Binary Preview Terms' license.
Claude improves a Riemann-zeta bound with 60 subagents and 31M tokens
Anthropic disclosed on August 10 that an unreleased research version of Claude improved the longstanding lower bound on the fraction of Riemann zeta zeros that satisfy the Riemann hypothesis from 41.6% to 67.2% by synthesizing recent papers rather than solving the hypothesis itself. Running inside Claude Code across two sessions, the model burned 31M output tokens, generated 650 initial ideas, then orchestrated ~60 subagents that ran 2,400 shell commands and thousands of numerical validation checks. The company frames it as a data point on the agent-orchestration approach to hard math problems.
OpenAI's new GPT-5.6-Cyber found two Chrome zero-days
OpenAI on Aug 10 expanded its Daybreak initiative with two tiers: Daybreak Blue (GPT-5.6 Sol with system-level cyber guardrails removed, which answers ~2% of advanced security queries) and Daybreak Red, which grants access to a new purpose-trained model, GPT-5.6-Cyber, that responds to 95% of sensitive queries covering exploit-chain development, authentication bypass and privilege escalation — up from 57.3% for its predecessor GPT-5.5-Cyber. The model discovered two previously unknown V8 vulnerabilities in Chrome that can be chained to corrupt memory and bypass the V8 heap sandbox; Google patched them under CVE-2026-15903. GPT-5.6-Cyber is OpenAI's first model to hit the 'High' cyber capability threshold under its Preparedness Framework (short of 'Critical', which paused Astra last week), and OpenAI is making hardware security keys mandatory for all Daybreak accounts on Sept 1.
bobdahacker.com
12h ago
ALERTA 30
AI notetaker tl;dv leaked 181K meetings, sat on the fix for six months
Security researcher bobdahacker publicly disclosed that tl;dv, a popular AI notetaker for Zoom, Google Meet, and Teams, left 181,874 meeting records across 84,312 users and 35,003 email domains queryable by any authenticated user due to a missing Firestore tenant-isolation rule. Roughly 1,000 records were public and 715 invitee emails exposed; the researcher reported the flaw January 28, 2026 and it remained unfixed through repeated follow-ups, with active recording sessions including government-agency, university, HubSpot, and Confluent calls joinable by outsiders.
Meta releases Muse Glimmer, a 30B agent model that runs on a laptop
Meta released Muse Glimmer, a 30B-parameter dense multimodal model under Apache 2.0, tuned for local agentic tool use, coding, and LLM-as-judge with a 131K context and support for 100+ languages. 4-bit quantization compresses it under 20GB so it runs on a single consumer GPU, hitting 3.1x speedup on RTX 5090 via speculative decoding. Meta paired the drop with a 6,500-word Zuckerberg essay promising open weights for Muse Spark 1.2 in the coming weeks, defending model distillation, and a $1B community fund for regions hosting Meta data centers.
portswigger.net
3d ago
ALERTA 30
AI research system finds novel HTTP desync attacks in 700 live sites
PortSwigger's James Kettle unveiled HTTP Terminator, an autonomous AI research system that tested 30,000 candidate desync vectors against thousands of authorized websites and identified roughly 700 vulnerable targets, including banks, government infrastructure, security products and an airport. The system generated new attack classes including a dual-matching Content-Length pattern, a 'dangling-byte' technique for more reliable response queue poisoning, and shared-parser confusion, and a human-guided cascade also exposed an Apache Traffic Server zero-day. Kettle frames it as the first case of AI producing genuinely novel security research, rather than reapplying known bug classes.
Cloudflare's Kitesurf agent browser uses up to 7x less memory than Chromium
Cloudflare launched Kitesurf, a cloud-hosted browser purpose-built for AI agents that runs inside Workers V8 isolates. Built in 12 weeks by stitching together Blitz (renderer), Firefox's Stylo (CSS), Parley (text) and Boa JS, Kitesurf passes ~215,000 Web Platform Tests and reports 3.1x-3.8x less CPU and 4.7x-7.0x less memory vs Chromium for screenshotting and HTML extraction. Available free in beta via Browser Run; the pitch is that agents don't need themes, tabs or extensions and would rather trade rendering fidelity for token-cost and context-window efficiency.
Rippling built an AI cost tracker after AI spend hit 40% of R&D budget
Rippling launched AI Spend Console after its own AI-token bill was on track to consume 40% of R&D headcount budget, growing 80% month-over-month with 10-15% of employees driving 60% of spend and one engineer burning $50K/month. The tool maps spend per employee and team against productivity signals (code output, PRs) and routes across Cursor, OpenAI, Anthropic, Grok and Z.ai's GLM 5.2 — which CEO Parker Conrad calls '85% cheaper but nearly identical performance.' Token spend dropped from 40% to 15% of headcount budget; July costs were 37% of April despite similar 600B token volumes.
Anthropic loosens Fable 5 on biology, cuts blocked queries by 85%
Anthropic rewrote and retrained Fable 5's biology safety classifier to distinguish everyday health, education and clinical questions from dual-use research. The company says the change cuts biology-related fallbacks by about 85% and total fallback volume by ~67% on Claude.ai, 55% on Cowork, 17% on Claude Code and 7% on the Claude Platform. Virology, toxicology and molecular-design prompts still route to Opus 5, so Anthropic warns Fable 5 remains 'not yet usable for professional biology research and drug development.'
Claude Code makes auto mode the default, and human reviewers look worse for it
Anthropic will flip Claude Code's Auto Mode on by default for Pro, Max and Team users starting August 14, replacing manual approval prompts with a classifier that vets each tool call for irreversible or destructive actions. Anthropic says internal testing across 1,000+ paid users showed the classifier caught 89% of dangerous commands compared to just 13.6% for human reviewers, and teams using auto mode ship roughly 25% more pull requests. The company will stop charging for the extra tokens the classifier consumes.
Anthropic lets Claude Code sessions message each other
Claude Code v2.1.224 introduces cross-session messaging: one session can now send a summary to another mid-task rather than forcing users to re-explain context. Claude composes the actual message from a user hint, so it's coordination rather than a raw history dump. Permission approvals and configuration changes are excluded, and any privileged actions still prompt the receiving session. macOS and Linux only for now.
Stanford's AI designs 16 novel bacteria-killing viruses
Stanford researchers published in Science on Aug 6 the design of 16 novel bacteriophages generated by their genomic AI model Evo, which was trained on ~2M viral genomes. The AI-designed viruses successfully infected E. coli — some killing bacteria faster than the natural ΦX174 template. The team excluded human-pathogen data, but biosecurity experts including Dr. Moritz Hanke warn the technique could lower barriers for bioweapon design; Imperial's Tom Ellis says built-in genetic-data restrictions could mitigate the risk.
OpenAI's first gadget is a $300+ doughnut speaker, Gurman says
Bloomberg's Mark Gurman reports OpenAI's first consumer device — designed with Jony Ive's LoveFrom — is a battery-powered, screenless smart speaker roughly the size of a hockey puck, doughnut-shaped, with a camera, microphones and mechanical parts that shift so it appears 'alive.' Pricing is pegged at $300-$400 with a targeted 2027 ship, though Apple's trade-secrets suit over metal finishing techniques could delay it. The pitch: a portable, humanlike ChatGPT companion for the home.
AMD buys Taalas, which etches AI models into silicon
AMD said Wednesday it acquired Taalas, a Toronto startup founded in 2023 that bakes AI model weights directly into custom silicon rather than storing them in HBM. Terms were not disclosed; the deal is expected to close in Q4 2026 subject to regulatory approval. Taalas's first test chip HC1, on TSMC's 6nm process, hit ~17,000 tokens/sec serving Llama 3.1 8B — a claim of roughly 48x Nvidia GPUs and 8.5x Cerebras at time of announcement — and a 20B-parameter HC2 is due this summer.
Mathematicians say OpenAI's proofs plagiarize prior work
Steven Miller (Yeshiva) says the sphere-packing proof OpenAI showcased pastes in the central argument from his own 2016 paper without credit, calling the pattern deliberate. Francesco Fournier-Facio (Cambridge) says the soficity 'breakthrough' stitches together ideas from 2016 and 2019 papers. Follows the July 30 OpenAI drop of 10 Astra proofs previously covered as a breakthrough.