Hundreds of contractors hired through Crossing Hurdles and paid via Mercor are reviewing real ChatGPT conversations under an internal effort called Project Lily, 404 Media reports. Workers score responses and can encounter prompts containing sensitive personal details. OpenAI says usernames are removed and personally identifying information is stripped, though details inside a conversation can remain. Human review is not unique to OpenAI, but the reporting makes its scale and privacy tradeoff unusually visible.
AI news for Tuesday, September 15, 2026
The Daily AI Espresso — the links the most-followed people in AI actually shared, curated every morning. Edited by Alexis · Live updates →
6 developments clear today's highest-consequence bar.
China’s Foreign Ministry used a regular press conference to rebuke Anthropic CEO Dario Amodei’s call for continued US restrictions on advanced AI chips and chipmaking equipment. Spokesperson Guo Jiakun called the essay fearmongering and said confrontation would disrupt global AI governance. The exchange turns a lab CEO’s policy argument into a direct government-level dispute over access to frontier compute.
Sam Altman told Fortune that OpenAI will not go public in 2026, calling the current moment ill-advised given safety concerns and saying the company still has work to do. He also floated the possibility of labs agreeing to pause when models reach new capability thresholds. That is an idea, not a commitment; the concrete development is that a 2026 listing is off the table.
A threat actor built exploits for two PaperCut NG/MF flaws, then used AI agents to automate the intrusions, according to GreyNoise reporting summarized by Help Net Security. The campaign reached at least 440 instances across 395 organizations in 48 countries; in one burst, 11 organizations were compromised in 26 seconds. The agents used an OpenAI Codex harness and DeepSeek, but the exploit development and targeting began with a human operator.
Australian AI data-center company Firmus is seeking to raise as much as A$7 billion—about US$5 billion—in an ASX listing targeted for the end of October, Data Center Dynamics reports. The size and timing can still change. Firmus raised US$2 billion in August at a reported US$10.5 billion valuation and is meeting investors across Asia; proceeds would support its compute buildout.
Sakana AI launched Fugu Max, an orchestration service that routes each request to the open model it judges sufficient. It costs $2 per million input tokens and $6 per million output tokens. Sakana claims its output price is 40% to 60% below several frontier alternatives and says it leads six benchmark categories; those comparisons are vendor-reported. A higher-capability Fugu Ultra v2 uses the same routing architecture.
We track thousands of AI experts across social media and YouTube to find what matters—not whatever wins the algorithm.
Turn the firehose into your briefing. Pick the topics, companies and people you track. Your agent sends a focused edition weekly—or sooner when enough important news breaks (up to three times per week).
The rest of today’s expert-filtered signal, ranked for consequence, utility, surprise, and range.
Apple opened the English public beta of its rebuilt Siri AI, powered by models custom-built with Google and Gemini. Processing is split between devices and Apple’s Private Cloud Compute. More languages are due next month, but the launch excludes the EU on several platforms and remains on hold in China. Apple also applies daily limits to some server-backed features, making this a consequential but deliberately constrained first release.
In pre-release testing on DeepSeek V4 Pro, SemiAnalysis measured Nvidia’s Rubin NVL72 at 59.4 million tokens per second per megawatt at 100 TPS, versus 28.5 million for GB300 Blackwell—a 2.1x gain. The advantage reached 7.2x at 150 TPS in the tested software configuration. The analysis was conducted with access to unreleased hardware and vendor assistance, so the numbers are directional, workload-specific and not yet independent production results.
Shanghai AI Laboratory introduced Atria Dawn Preview in a 143-author preprint and evaluated it across 16 benchmarks. The paper also analyzes 769 task records from 56 people who used the model during its development; participants rated about one-third of completed AI-assisted tasks infeasible without AI. That is a participant judgment, not proof of independent autonomy, and the authors stress that humans retained final decision authority.
Intel spinoff Cornelis Networks raised $205 million and introduced Active Compute Fabric, an open, GPU-agnostic networking layer for AI clusters. Its current 400 Gbps CN5000 switch is shipping, with an 800 Gbps CN6000 generation planned next. The pitch is to give data-center operators another interconnect path beyond Nvidia’s tightly integrated stack; product-performance claims still need independent deployment evidence.
Altman, Amodei, Musk and Hassabis backed a call to pace frontier development, presenting it as a safety response. A more cynical interpretation is spreading online: frontier labs are burning enormous sums on compute, and slowing together could conserve cash and protect incumbents from cheaper or open rivals. Axios notes that pacing could reduce Anthropic’s compute spend—but also that no frontier lab can simply sit out the capital markets. Reddit commenters go further, saying the labs are running out of money.
Useful releases, methods, and workflows worth trying.
Perplexity’s Portable Computer agent is now available in its Windows app for Nvidia GeForce RTX and RTX PRO GPUs with at least 24GB of VRAM. Nvidia says the model, agent harness, orchestrator and scheduler run on the device, defaulting to a locally optimized Qwen model. Users can choose cloud models, but the app asks permission before sending a task off-device. The hardware floor keeps this local-agent launch aimed at high-end PCs.
Anthropic says its continuous-integration workload grew 25-fold in six months as engineers shipped roughly eight times more code per quarter than during 2021–25, with Claude authoring about 80% of it. Test volume rose 10-fold. The company responded with test-impact analysis that selects checks based on code changes. These are self-reported internal figures, but they reveal an important second-order cost of agentic coding: verification infrastructure has to scale too.
What readers are testing in the wild—reported with its limits, not mistaken for a benchmark.
A Reddit user connected an ESP32, a camera and a tiny cat-shaped thermal printer to Codex, then let the agent work through the terminal and images. Codex wrote and flashed firmware, built a web interface and used photographs to check its own output; the final print was a picture of the printer itself. The user handled physical checks, and this is a self-reported build rather than a benchmark—but it is a delightful glimpse of agents crossing from code into hardware.
That’s today’s shot. — Alexis · AI Weekly
