The UN-backed Independent International Scientific Panel on AI issued its first thematic brief on September 21, warning that safeguards for increasingly capable agents are not keeping pace. The brief uses a May-to-July OpenAI and Hugging Face test as a case study and discusses incident reporting, independent scrutiny, and layered controls; it is a warning and policy input, not a binding UN rule.
AI news for Tuesday, September 22, 2026
The Daily AI Espresso — the links the most-followed people in AI actually shared, curated every morning. Edited by Alexis · Live updates →
14 developments clear today's highest-consequence bar.
Grok 4.7, xAI's newest coding and knowledge-work model, ships at $2 per million input tokens and $6 per million output tokens, with a fast variant at double the price and double the output speed. Benchmarks published by the company put the model at 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1, 64.0% on EEBench, 56.7% on HealthBench Professional and 19.6% on the Harvey Legal Agent Benchmark. "Grok 4.7 is our most capable model for coding and knowledge work," the release states.
In a Truth Social post Friday, President Trump said he will appoint an AI czar and stand up a new 'AI Force' modeled on the Space Force to oversee the industry, adding that 'only High I.Q. individuals need apply.' He dismissed AI safety concerns as a 'hoax' and vowed the White House 'will not in any way hinder or stifle' AI growth, framing US leadership as necessary to 'beat China.' No appointee was named.
Plugin4Shell, disclosed by researchers at AIR and covered this week in Help Net Security, lets an attacker swap malicious code into a pinned plugin on four major AI coding agents without any user action. AIR calls it "the first supply chain vulnerability of the AI agent ecosystem" and reports that "925 skills already in active use had been hijacked," reaching "134,000 agents." The bug breaks SHA pinning, the mechanism that is supposed to lock a plugin to a specific reviewed commit.
Alibaba DAMO Academy released Damo Radar, a vision-language model trained on 420,000+ contrast-enhanced abdominal CT exams and 15M anatomy-focused image-text pairs, with weights, code and training framework on GitHub and Hugging Face. On ~40,000 real-world exams it averaged AUC 0.913 across 146 clinical findings and outperformed 23 of 26 expert radiologists in a head-to-head study published in Science.
A three-person team at security startup Hacktron AI chained two vulnerabilities that started with a HEIF image upload to OpenAI's community forum and ended with logged-in access to OpenAI employee ChatGPT and Codex accounts. The full exploit chain came together in under 72 hours. OpenAI paid the team $6,500 for the report, TechCrunch's Aditya Mehta and Rebecca Bellan wrote. The initial entry point was community.openai.com, which runs on Discourse. The forum decodes uploaded HEIC and HEIF images through ImageMagick and libheif. A heap buffer overflow in libheif let a specially crafted image corrupt server memory.
Z.ai's ZCode coding agent silently packaged 42,411 files from one developer's workspace into a 313MB encrypted archive and made 564 failed attempts to ship it to Alibaba Cloud, according to Tom's Hardware. The finding came from a researcher publishing as ferstar. In a reverse-engineering write-up dated September 18, ferstar traced the client requesting upload credentials from zcode.z.ai, receiving an RSA public key and Alibaba Cloud object-storage form signatures, and encrypting the payload with AES-256-CTR before shipping it.
Anthropic has a wet biology lab in the San Francisco Bay Area, and its head of life sciences confirmed it this week to TechCrunch. "We believe that to do biology, the final test is still, and will be for a while, in real lab work," said Eric Kauderer-Abrams, Anthropic's head of life sciences. "We absolutely are doing that today." The company has not disclosed the lab's size, staff count, opening date, or biosafety level. Kauderer-Abrams said the focus is fundamental biology rather than drug discovery, and that some biological work still goes to outside partners.
PrismML's Ternary Bonsai 2 27B takes Alibaba's Qwen3.8 27B and squeezes it from 53.80 GB down to 5.93 GB by holding every language-model weight to one of three values: minus one, zero, or plus one. According to MarkTechPost, the ternary scheme lands the model at roughly 1.72 bits per weight, with only 26.2 million parameters, or 0.0976% of the total, kept in higher precision for recurrent state paths and normalization. "Each group of 128 weights shares 1 FP16 scale," the writeup notes, and a blockwise Hadamard rotation inspired by SpinQuant is applied before ternary assignment.
Google's open-source agent orchestrator AX (Agent Executor) reached v0.3.0 and took the top AI slot on Hacker News with 481 points. The release splits AX into three services — an API frontend, a reconciler, and a sandboxed task runner — and moves task state out of Kubernetes custom resources into Redis Streams because etcd was not built for the churn of millions of short-lived agent tasks. AX runs on top of Agent Substrate and is Apache-2.0 licensed.
Alibaba's Qwen team released Qwen3.8-Omni-Flash on Sept 18, a native omnimodal model that jointly processes text, images, audio and video with a 1M-token context. On roughly 30 evaluations it beats Qwen3.5-Omni-Plus by 26%+ on average, with a 45.7% token-usage reduction on agentic video tasks. API-only via QwenCloud, Alibaba Cloud Model Studio and Qwen Studio at $0.15/M input and $0.47/M output; no open weights at launch.
Alibaba's Qwen team pushed Qwen-Image-2.1 to Hugging Face and ModelScope today, pairing a 7B, 32-layer single-stream DiT with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE that outputs native 2048x2048 at 40 steps. The release ships two 9B PE-T2I/PE-I2I prompt-rewriter checkpoints, supports up to 10 reference images and mask/circle-based local edits, and switches licensing from Apache 2.0 on the earlier Qwen-Image line to a non-commercial Qwen Research License Agreement — commercial users now need a separate agreement.
StepFun officially announced Step 5 Preview on September 20, a 600B-parameter sparse MoE with 27B active per token and a 1M-token context, and opened API access the same day. Artificial Analysis pegs the model at 44 on its Intelligence Index — matching Kimi K3 Max and roughly a seventh the price of GPT-5.6 Sol — at $1 per million input tokens and $2.70 per million output tokens with a 95% cache discount. StepFun says full open weights follow on October 15; the Hugging Face repo currently ships only a .gitattributes file.
Sony Music and Universal Music Group filed a 45-page complaint against Suno on Friday in U.S. District Court in Massachusetts, arguing that Suno's new 'v6' music model is still built on copyright-infringing training data. The complaint's central metaphor is blunt. 'V6 is not a fresh start; it is the fruit of the same poisoned tree,' the labels write, according to Variety.
Pick the topics, companies and people you track. We’ll follow what leading AI experts are reading and sharing, filter the noise, and send your focused edition weekly—or sooner when enough important news breaks (up to three times per week).
Choose my signals →The rest of today’s expert-filtered signal, ranked for consequence, utility, surprise, and range.
OpenAI says an internal model has resolved more than 100 additional open problems and has created a nine-member Advisory Group on Mathematics and Artificial Intelligence at Princeton's Institute for Advanced Study. The group can assess results and coordinate releases, but TechCrunch reports it cannot slow or redirect OpenAI's internal research, and the Institute says final decisions remain with the company.
Xiaomi released MiMo-V2.6-Pro-RL weights, code, and a technical report under an MIT license on September 21. The 1.02-trillion-parameter mixture-of-experts model activates about 42 billion parameters per token and scored 46 on Artificial Analysis' Intelligence Index, tying Grok 4.7 there. Xiaomi reports roughly $2.62 million for its large reinforcement-learning run; that cost is vendor-reported, not an independent audit.
The CARE preprint collects failed robot-policy rollouts, turns representative failure states into corrective demonstrations, and uses 3D monitoring to trigger targeted fixes without restarting the whole task. Across several vision-language-action backbones, the authors report average task-success gains of 14.5 points in simulation and 15.9 points on real-world dual-arm tasks; these are author-reported preprint results.
The onPanda preprint replaces full-response editing with a locate-correct-continue loop: an annotator fixes the first unsuitable token, then lets the model regenerate from that corrected prefix. The authors report a 52% reduction in median annotation time and say the approach captures token-level preference signals for later alignment training. The result is from a new preprint and has not yet been independently replicated.
Amazon says it has blocked Meta's Muse agent from shopping on Amazon.com after Meta declined to exclude the site. Amazon argues the agent does not identify itself and may handle account credentials without Amazon's consent; Meta says Muse cannot see passwords or payment methods because credentials enter secure storage. The clash tests who controls the customer relationship when agents transact across services.
Johns Hopkins researchers added hedges, collective phrasing, and expressive adjectives associated with women's American-English usage to workplace prompts sent to GPT-4, Llama, Gemma, and Mistral. Across all four, the replies became shorter, less formal, and less complex even after the team accounted for tone; changing only the signer's name had little effect. The work is scheduled for the October Conference on Language Modeling.
Useful releases, methods, and workflows worth trying.
Linear says agents helped its test suite nearly quadruple this year, turning continuous integration into the new bottleneck. By changing runners and compilers, shrinking critical-path checkouts, batching short checks, and selectively sharing test state, it kept pull-request waits near five minutes while roughly halving runner time per test. One batching change alone saved about 87,000 runner-minutes a month.
What tracked AI experts are sharing on social platforms—verified before it reaches your inbox.
A September 22 Reddit post showed photorealistic street-art concepts generated in ChatGPT. The poster said the visual ideas were largely theirs while the model supplied variations; replies split over whether prompting counts as making art and whether 'street art' still means anything when it never reaches a street. The images are playful, but the argument is a live test of authorship and authenticity as generation gets harder to spot.
That’s today’s shot. — Alexis · AI Weekly
