Anthropic's Claude Opus 5.5, released September 22, lists at $4 per million input tokens and $20 per million output. That is 20% below Opus 5. Cache reads drop 60% to $0.20 per million, which the company notes make up the majority of agentic and coding work costs. Anthropic puts the overall savings at roughly 40% on typical workloads. On its own benchmarks, Opus 5.5 posts 66.4% on Terminal-Bench 4.0 (up from Opus 5's 52.3%) and 81.8% on OSWorld 2.0 for computer use.
AI news for Wednesday, September 23, 2026
The Daily AI Espresso — the links the most-followed people in AI actually shared, curated every morning. Edited by Alexis · Live updates →
4 developments clear today's highest-consequence bar.
OpenAI has fired multiple contractors hired to grade ChatGPT's answers after catching them using AI to do the work, 404 Media reported. Internal documents reviewed by the outlet reference "more than ten thousand contractors" across the company's rating pipelines and bar them from touching AI-detection tools, GPTZero, Grammarly, or a chatbot itself while working.
OpenAI launched GPT-6 Sol and Luna on Tuesday, priced at half the GPT-5.6 tier ($2/$10 per million tokens for Sol vs. $4/$20 previously) and positioned as 'cut from the same cloth' as flagship Astra. Sol is aimed at complex coding while Luna targets high-volume clerical work; OpenAI says Sol makes about half as many mistakes as its predecessor on internal factuality evals, reaching Astra-level reliability. The drop landed roughly 90 minutes after Anthropic released Claude Opus 5.5, with OpenAI claiming both models beat Fable and Opus on many tasks.
Alphabet's Intrinsic Innovation released Intrinsic Core, a Robot Operating System-compatible robotics stack, under an Apache 2.0 license at ROSCon 2026 in Toronto, SiliconANGLE reported. The bundle includes Intrinsic Control, a hardware-agnostic real-time control framework; pose estimation built on Nvidia's FoundationPose; motion planning; grasp planning; simulation and calibration services; and Intrinsic-ROS drivers for sensor and third-party hardware integration.
Pick the topics, companies and people you track. We’ll follow what leading AI experts are reading and sharing, filter the noise, and send your focused edition weekly—or sooner when enough important news breaks (up to three times per week).
Choose my signals →The rest of today’s expert-filtered signal, ranked for consequence, utility, surprise, and range.
404 Media reports that Meta publicly marketed Muse's ability to make phone calls to businesses on customers' behalf, but internal communications show many calls were routed to a 'human agent layer' in a call center. Internal testers say they were not told the caller was human until after the call concluded, and one employee warned 'this has potential for so much negative PR' by portraying Muse as 'not good enough' without human backup.
"Claims of incipient, dangerous superintelligence are not based in good scientific or engineering practice," Timnit Gebru and Emily M. Bender write in MIT Technology Review. Gebru, executive director of the Distributed AI Research Institute, and Bender, a linguistics professor at the University of Washington, argue that the summer's breathless AI announcements collapse once experts actually look at them. Take the OpenAI–Hugging Face incident.
Research: Argument that AI firms lack reliability in self-regulation; regulation, corporate responsibility, technology, trust
article: analysis of AI-driven data centers' energy demands and local labor responses in Pennsylvania's industrial landscape; data science, energy, labor, AI
AIDE², an AI research agent that rewrites its own code, cut its reward-hacking rate from 55% to 32% over an eight-day autonomous run, a property the paper says the loop never explicitly optimized for. The arXiv preprint from Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu and Zhengyao Jiang describes a loop in which the agent proposes changes to its own code, benchmarks modified versions of itself on AI R&D tasks, and keeps whichever performs best on hidden evaluations.
Business Insider profiles 23-year-old Noah Shinn, whose personal-agent startup Instinct soft-launched as an invite-only site distributed through VCs and is reportedly fundraising ~$1B at a $10B valuation with Sequoia and Benchmark — up from a $2.5B mark just weeks earlier. The agent, accessible via text or phone, has topped 100,000 users and hit capacity constraints as demand outpaces compute.
Alibaba Cloud will open its first regions in Turkey, Finland and the Netherlands over the next 12 months, with the Netherlands site first to come online in October, Bloomberg reported from the company's Apsara conference in Hangzhou. Alibaba executive Li Feifei delivered the announcement, which also covered footprint expansions in Malaysia, Germany, the United Arab Emirates, France and Hong Kong. Combined with the three new regions, that adds up to eight locations across the coming year, DatacenterDynamics wrote.
What tracked AI experts are sharing on social platforms—verified before it reaches your inbox.
StableVQ, a tokenizer-training method from a team at Huazhong University of Science and Technology, KlingAI Research and South China Normal University, reports full codebook utilization with rFID 1.05 on ImageNet at a 262,144-code book, according to the paper on Hugging Face. The authors trace Vector Quantization training failure to the "entanglement of the Encoder–Decoder and Codebook training." Their remedy is three parameter-free changes: a Dynamic Straight-Through Estimator that reweights encoder gradients by quantization error, a Region VQ Loss that propagates targets from active codes to inactive ones, and…
That’s today’s shot. — Alexis · AI Weekly
