GPT-6 Astra just went live, and everyone is testing its limits

What's trending in AI right now, from the app charts to the community feeds. Real links, our take.

GPT-6 Astra just went live, and everyone is testing its limits

OpenAI launched GPT-6 Astra this week, first as a limited preview on September 3 and then more broadly on September 5. It's already the top conversation everywhere: r/ChatGPT is full of people reporting that Astra burns through tokens extremely fast, that it asked one user if it could log into websites on their behalf, and that a pre-release version partially solved FrontierMath's new Erdős-level benchmark. 9to5Mac has a solid rundown of what changed, and The New Stack covers the benchmark numbers. If you're on a Plus plan and planning to try it today, expect to hit your monthly allocation faster than usual.

Google launched Gemini 3.8 Flash in the same 72-hour window

A few days before Astra landed, Google made Gemini 3.8 Flash generally available. It has a one-million-token context window and scores well on software engineering benchmarks against models that are significantly larger. The Google AI blog has the full details, and The Register calls it Google reminding everyone it's still in the race. Two of the biggest AI labs shipped flagship updates in the same week, which is becoming the normal pace of this year.

Conduit jumped 17 places in the App Store this week

The biggest mover in the productivity charts right now is Conduit, an open-source mobile client for self-hosted AI models and Open WebUI instances. It climbed 17 spots to reach number five in productivity. The reason people are gravitating to it is straightforward: your chats go directly to whatever AI backend you point it at, with no third-party server in the middle. It's on the App Store and the source code is on GitHub. If you already self-host Ollama or run Open WebUI, this fills a real gap.

Plaud keeps climbing in productivity

Plaud, the AI-powered meeting recorder, has been rising steadily in the charts this week. The company makes hardware devices that record conversations and a companion app that transcribes in 112 languages and generates AI summaries. They report 2.5 million users globally now. The official site walks through the full hardware lineup. The category of physical AI recording devices is clearly finding its audience, and Plaud is the one moving in the charts right now.

Qwen built a benchmark where AI models run fake online stores for a year

A genuinely novel benchmark came out this week from the Qwen team. Called E-Commerce Bench, it gave 18 different frontier AI models a virtual budget and tasked each with running a simulated online store for 365 days using real product data, with fraudulent suppliers mixed in among the legitimate ones. The goal is to test how AI agents handle hundreds of sequential decisions over a long time horizon, rather than a single-shot question. The benchmark page has the full methodology. It's an early look at how differently today's models perform when they have to keep managing something over time, and the results across models are pretty spread out.