Or OpenAI's blog https://t.co/rnSrr68wy0 And attend today's Jalapeño Hot Chips session with Richard, Ravi, and Chris from OpenAI too. (7/7)
SemiAnalysis
Tracked through public AI activity and peer connections inside the directory.
- AI signals
- 10 past 30d
- Sources
- 4 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 6d ago
Articles & links
Read more at our article👇️ (6/7) https://t.co/dbuAUrRSt1
Gemini is Cooked but GCP is Cooking GCP YoY rev growth >100%, DeepMind's long term failure is Google Cloud's short term gain https://t.co/PqVUCF8qSg
Are Open Models Catching Up? Comparing open vs. closed models across the eras of frontier models, Is the gap narrowing? https://t.co/TBZnvquwio
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? $3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200 https://t.co/PgtKNYXdFr
- SemiAnalysis's AgentX 1.0, built on 393 anonymized Claude Code traces at 1M+ context, cost more than $3M and used ~2MW across 1000+ chips.
- On Qwen3.5 SGLang the report puts Nvidia at 'over 20x better performance' at 90 tok/s/user; B300 FP4 shows '12x better performance per dollar' vs H100.
- AMD's ATOM stack shows strong single-GPU kernels but almost no production adoption — only one Alibaba ad unit runs it live, the authors say.
TokenBudgeting: Our Conversations with Enterprises on Token Spend Was Widespread TokenMaxxing Ever Really Here? https://t.co/tgNssBBGVk
- Meta employees consumed over 60 trillion tokens in a 30-day window in early 2026, with one individual alone accounting for about 280 billion.
- Monthly per-employee caps now range from $250 at an aerospace and defense manufacturer to $2,000 at Workday and Stripe, with no cross-industry consensus.
- Ramp data cited by SemiAnalysis shows 99th percentile customers spend about $90,000 per employee per year while the median customer spends $136.
https://t.co/kmFqrIAzF2 (2/2)
- AgentX v1.0 turns 393 opt-in Claude Code sessions into a replay benchmark, with a median 142k input tokens and 444 output tokens per request.
- To share traces safely, inputs are reduced to session-scoped chained hashes in 64-token blocks that preserve matching prefixes but not content.
- SemiAnalysis says the benchmark's biggest first-months output was 50+ upstream pull requests from partners including vLLM, SGLang and TensorRT-LLM.
Cerebras's Next Generation CS-4: Fast Just Got Faster Double the Performance, Double the Power, Double the Fun https://t.co/qUgaBrNOYA
- Cerebras CS-4 delivers more than 4,400 tokens per second per user on GPT-OSS-120B, up to 30 times a GPU inference baseline.
- Each rack pairs three WSE-3 Turbo wafers for 750 PFLOPS, 129.6 PB/s of memory bandwidth, and two-microsecond wafer-to-wafer latency.
- AMD Instinct and AWS Trainium handle prompt processing while CS-4 acts as the decode accelerator; first shipments begin this quarter.
We discuss the potential TPU beneficiaries in our Core Research Exclusive Note, "TPU Long and Prosper: Don't Cry Because Gemini's Over, Smile Because TPU Happened." (3/3) https://t.co/Onejdf8iMG
- SemiAnalysis projects more than $100 million in 2026 revenue, up from roughly $20 million a year earlier, according to The Information.
- Core Research is one of four institutional products, aimed at hedge funds, long-only asset managers, venture investors and corporate strategy teams.
- The Substack newsletter reaches 200,000-plus subscribers but is sold separately from the institutional models and Core Research.
To learn more about ICI networking topology, please refer to our networking model (8/8) https://t.co/kqt0IiSiFi
- SemiAnalysis is offering device-level tracking of AI cluster networking across five fabric layers, with data running 2023 to 2026.
- Coverage spans 80+ hyperscaler configuration panels for Microsoft, Google, Meta, Amazon, Oracle, X.AI and neoclouds, tied to specific accelerator SKUs.
- 25+ suppliers are tracked, including Nvidia, Arista, Broadcom, Cisco, Coherent and Lumentum, across 200G to 1.6T transceiver speeds.
To learn more, check out our Datacenter Model, which covers BitDeer, and our AI TCO Model, which covers the economics of NeoClouds. 6/6🧵 https://t.co/wACjpkQLM8 https://t.co/EeBYF4i2ve
- SemiAnalysis's AI Cloud TCO Model covers Nvidia, AMD, Intel, and custom accelerators across cost of ownership per hour, inference cost per million tokens, and training cost per FLOP.
- The $/hr calculation is built from upfront server capex, system power consumption, colocation and electricity costs, and cost of capital, wrapped in a three-statement financial model.
- Detailed install base projections run through 2028 and vendor unit shipment estimates through 2034, aimed at operators, procurement teams, and equity and debt investors.
Meta Compute: Everyone Wants To Be A Cloud Zuck Takes Plan B? SpaceX 2.0, Bedrock 2.0, MSL Isn't Giving Up, Scaling RecSys by 10x... ClusterMAX ranking coming soon? https://t.co/RKHoW1Du89
- Meta has contracted over 5GW of capacity across cloud and colocation, per SemiAnalysis, after nearly 10GW of deals since early 2024.
- Meta is reportedly in final talks with Anthropic to host private Claude instances, akin to Bedrock or Vertex from rival hyperscalers.
- SemiAnalysis pegs Meta's strategy as four-track: frontier training, 10x-plus ads recsys, Claude hosting, and SpaceX-style external rentals.