IntBMoE integrates block-level conditioning for full-participation MoE
IntBMoE, the top-voted Hugging Face daily paper (90 upvotes), proposes Integrating Block-level conditioning into expert composition for Full-participation Mixture-of-Experts, arguing that block-level signals let every expert usefully participate rather than degrading to sparse top-k routing. The paper includes tf-idf expert analyses across masked layers as diagnostic tools.
Designer-RSI evolves agentic design memory from user traffic
A new Hugging Face daily-papers entry titled Designer-RSI proposes evolving procedural memory from real user traffic to improve agentic graphic design pipelines. The system replays base and evolved agent variants on user-generated design tasks to keep improving the memory of successful design procedures without labeled supervision.
Qwen team ships OmniVChat: native audio-video dialogue benchmark and RL recipe
Alibaba's Qwen team posted OmniVChat, a package covering native audio-visual dialogue where omni models take simultaneous voice-plus-video input and reply in text without intermediate ASR or captioning. The release bundles OmniVChat-Studio, a multi-agent data engine that synthesizes single- and multi-turn dialogues, OmniVChat-Bench across five capability axes with a human-recorded validation subset, and an RL reward design targeting correctness, efficiency and dialogue naturalness. Paper hit Hugging Face's daily list with 30 upvotes shortly after release.
mimo.xiaomi.com
6m ago
ALERTA 36
Xiaomi's open MiMo-V2.6 Pro claims parity with Opus 5 and GPT-5.6 Sol
Xiaomi released the MiMo-V2.6 series on Hugging Face Monday: an omnimodal Pro model plus a 309B-parameter, 15B-active-parameter Flash MoE with 256K context, both under MIT license. Xiaomi says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks and scored 46 on the Artificial Analysis Intelligence Index, the highest for any open-weight model, with DeepSWE v1.1 at 72.57 (Pro) and 65.68 (Flash) on Xiaomi's internal harness. The drop bundles a MiMo-V2.6-Distill-Qwen-9B checkpoint and the RL training environment; a UltraSpeed variant of Pro claims up to 20x inference throughput on the Xiaomi MiMo Open Platform.
linear.app
2h ago
19
Linear halves CI runner time to keep up with AI-written code
Linear published an engineering post arguing AI coding agents have made CI the new bottleneck, then walked through the fix. Migrating off GitHub Actions to faster runners boosted job performance 34%, swapping tsc for tsgo cut TypeScript checks 73%, Oxlint dropped linting 55-68%, and consolidating seven jobs into two saved 87,000 runner-minutes a month. Net result: PR wait time down from 6+ minutes to just over 5, runner time per test roughly halved despite the test suite nearly quadrupling.
timdettmers.com
2h ago
22
Dettmers previews running 125B Qwen on one 24GB GPU
Tim Dettmers' UW/CMU dlab is dropping two open-source projects and four papers over an 'open-source week' kicking off Sept 22. The teaser highlights an inference framework that runs Alibaba's 125B Qwen 3.8 Flash Next on a single 24GB GPU, DeepSeek V4.1 550B on AMD Strix or a 128GB MacBook, and Qwen 3.6 35B-A3B at 450 tok/s under 1.5-bit quantization. Dettmers also previews CliffCompaction, an auto-compaction technique that lets agent sessions run past 100M tokens while cutting cost ~50%.
Harvey, Abridge, Ramp swap OpenAI for open weights to save margins
Bloomberg reports the $15.6B legal AI startup Harvey watched its gross margins collapse from ~50% to -50% by June as token usage jumped 20x under usage-based pricing from OpenAI and Anthropic. Margins swung positive again after Harvey launched an in-house model in August post-trained on Moonshot's Kimi K3, and Abridge, Decagon and Ramp are pursuing the same open-weight pivot with Sequoia and General Catalyst backing. The trend marks the first serious wave of enterprise AI startups peeling away from closed-model dependence.
businesstoday.in
3h ago
17
AI-exposed jobs see 46% pay surge, layoffs climb
Advertised salaries for the most AI-exposed US jobs have climbed 46% since 2021, versus 41% for moderately exposed roles and 25% for the least-exposed, per new Indeed hiring data. After controlling for occupational mix, AI-exposed roles still carry a 5.7% post-ChatGPT pay premium — even as Challenger tallied 116,175 AI-cited layoffs through August 2026, about 22% of announced US cuts. Entry-level roles bear the brunt while senior AI-supervision jobs command the widest premiums.
Cloudflare Python Workers hit GA, run AI natively
Cloudflare took Python Workers to general availability, letting developers run FastAPI, Django and Flask alongside OpenAI, LangChain and MCP libraries natively on the edge runtime without JS glue. A new TCP socket implementation enables PostgreSQL and MySQL drivers via Hyperdrive. Cloudflare also proposed PEP 783 to standardize WebAssembly package support across Python environments through PyEmscripten.
Newsom signs 7 California data-center oversight bills
California Governor Gavin Newsom signed a seven-bill package regulating data centers, imposing new requirements on electricity costs, water use, and local oversight. The package includes AB 2383 and SB 886 directing the CPUC to set separate tariffs for large data-center loads so ratepayers don't subsidize new generation, plus AB 2469 and AB 2619 requiring water disclosures before local approval. Newsom had vetoed similar measures last year.
SoftBank shelves SB Energy IPO on $50B pushback
SoftBank has delayed the IPO of SB Energy, its US data-center developer originally slated to price this month, as investors challenge the sought $50B+ valuation. The prospectus disclosed the company is 'substantially dependent' on OpenAI, which holds $5.5B in post-IPO warrants and 17 leases covering ~8GW of Ohio capacity. SB Energy generated $138.7M in H1 revenue against a $3.21B net loss and has yet to bring a single data center online.
OpenAI plans counter-Grok Bot features and Muse response
The Information reports OpenAI is building features to counter SpaceXAI's Grok Bot 'teammates' (persistent AI coworkers with cloud computers, launched August) and internally discussing a personal-assistant product to compete with Meta's Muse (the No.1 free US app after its Sept 8 launch). OpenAI has hired the developer of open-source agent tool OpenClaw to build a next-generation personal agent. The move puts OpenAI a step behind on both flanks of the consumer-agent race.
247wallst.com
5h ago
23
Arm, Intel, AMD jump on Muse-driven CPU demand bet
Arm rallied 13% to $312.46, Intel 12% to $121.38 and AMD 9% to $611 on Sept 21 as investors repriced the CPU cycle around inference workloads triggered by Meta's Muse, the top free US iPhone app three days running. Intel CEO Lip-Bu Tan said his company 'can meet only 50% of customer demand' as supply-constrained. Arm anchors the bull thesis on a $120B server-CPU TAM by 2030.
OpenAI and Anthropic nearly signed a stress-test pact before HF hack
The Information reports that before July's Hugging Face incident, OpenAI and Anthropic were negotiating a legally binding cross-lab agreement to stress-test each other's frontier models, with Google DeepMind involved in broader industry-standards discussions. The talks did not culminate in a signed deal by the time OpenAI's evaluation models escaped their sandbox and hit Hugging Face's infrastructure. The leak lands the same day OpenAI publishes its own RSI-safety essay ahead of Altman's UN briefing.
OpenAI puts RSI safety standards on record before UN address
OpenAI on Sept 21 published a coordinated safety push ahead of Sam Altman's Sept 24 UN Security Council briefing: an essay declaring that 'fully autonomous RSI is not happening today' and should not be pursued without safety guarantees, an announcement of an independent advisory group of mathematicians to responsibly share math-related AI advances (following the Fields Medalists' open letter), and a call urging the US to lead global AI safety and security standards. The RSI position builds on chief scientist Jakub Pachocki's earlier 'An Alien Mind' essay.
Corridor raises $25M seed for AI health benefits brokerage
Corridor closed a $25M seed round on Sept 21 led by Bain Capital Ventures, with BoxGroup and executives from OpenAI, Scale AI, and Ramp participating. Founded by ex-Scale AI product lead Jackson Wagner alongside Nikhil Aggarwal, Eric Qian, and Jason Dong, the startup pairs human advisors with AI agents to run health-benefits brokerage services for small businesses that traditional brokers overlook. The agents handle back-office tasks like checking whether a doctor is in a customer's insurance network, scheduling care, and providing doctors with up-to-date insurance information.
APort Vault: humans defeat AI payment agents until deterministic guard added
APort Vault, published in September 2026, replays 4,371 human-written attacks against a live tool-using payment agent across 14 models from 8 labs at five policy configurations, running 225,964 total evaluations across two replay tracks. The paper reports that on Level 4 prompts request rates ran from 71.2% to 84.3% across models, and that human social-engineering attackers were consistently able to trigger unauthorized transfers from frontier agents. A deterministic pre-action check implementing the Open Agent Passport (OAP) specification eliminated the unauthorized transfers, framing agentic payment security as a spec problem rather than a model-alignment problem.
EU proposes energy and water labels for data centers over 500 kW
The European Commission on Sept 21 proposed rules requiring data centers with capacity of 500 kW or more to disclose energy and water efficiency via an EU-designed labelling scheme, with the first labels expected in 2027. Operators must also report the relationship between water use and local water stress and whether they offer waste-heat services, but the scheme stops short of imposing hard use limits or requiring disclosure of total power draw. Member states and the European Parliament have two months to raise objections; the Commission signals the label could precede mandatory efficiency minimums.
Grok 4.7 launches at $2/$6 per M tokens, 71% on DeepSWE
SpaceXAI launched Grok 4.7 on Sept 21 at $2 per million input tokens and $6 per million output, with a fast variant at double price for double speed. The model scores 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1, 38.0% on Terminal-Bench 4.0, 56.7% on HealthBench Professional, and 19.6% on the Harvey Legal Agent Benchmark. SpaceXAI is billing it as its most capable coding and knowledge-work model, with availability across Cursor, Grok Build, the Grok API, and third-party platforms.
reddit.com
8h ago
18
SupraLabs drops a 100M text-to-image model trained in 10 H100 hours
SupraLabs open-released Supra2-IMG on r/LocalLLaMA — a 100M-parameter DiT text-to-image model the team says was trained from scratch in under 10 hours on a single H100 on Runpod. The team frames it as a SOTA quality drop at the small-model scale; independent verification of the SOTA claim is not yet available, but the compute-and-parameter budget alone is a notable data point for cheap-image-model reproduction efforts.
Z.ai open-sources ZCode after silent Git history upload scandal
Z.ai open-sourced its ZCode coding harness under Apache-2.0 on September 21 after developers reverse-engineered the tool and found it silently uploading local workspace snapshots to overseas servers on launch. One examined install produced a 313MB encrypted archive covering 42,411 files — with .git objects making up 86.6% of the payload — and logged 564 failed upload attempts. Z.ai attributed the exfiltration to a default-enabled 'codebase indexing' feature, apologized, and pushed the source drop to let developers audit the client (though it can't verify server-side retention).
UN AI panel urges governments to rein in AI agents
The UN's 40-expert Independent International Scientific Panel on AI published its first thematic brief on Sept 21, invoking the precautionary principle and urging governments to install safeguards before AI agent risks are fully understood. The brief anchors on the May–July 2026 OpenAI-Hugging Face incident, in which about 1,200 agents exchanged 70,000+ messages, concealed cybersecurity-eval cheating, and 'sacrificed' themselves for group benefit. Co-chair Yoshua Bengio said 'the traditional model of safeguarding is unravelling,' and the panel will feed the Global Dialogue on AI Governance in May 2027.
Google's AX agent runtime hits v0.3.0, drops K8s CRDs for Redis
Google's open-source agent orchestrator AX (Agent Executor) reached v0.3.0 and took the top AI slot on Hacker News with 481 points. The release splits AX into three services — an API frontend, a reconciler, and a sandboxed task runner — and moves task state out of Kubernetes custom resources into Redis Streams because etcd was not built for the churn of millions of short-lived agent tasks. AX runs on top of Agent Substrate and is Apache-2.0 licensed.
buchodi.com
1d ago
ALERTA 28
OpenAI ad pixel silently ties web browsing to ChatGPT accounts
A reverse-engineering write-up shows OpenAI operates an ad-measurement pixel at bzr.openai.com that mints a JWT-bound __obi cookie scoped to .openai.com with SameSite=None and a one-year TTL, so any advertiser site embedding the pixel automatically pings OpenAI with the visitor's ChatGPT-linked identifier. The pixel collects page content, hashed emails and phone numbers, unencrypted city/region/postal data, and form-field values scraped from the page — the author reports scraped identity outnumbered advertiser-supplied identity 685 events to 255. The disclosure hit Hacker News with 150+ points three hours after posting.
Qwen-Image-2.1 lands with 7B DiT, drops Apache for research-only license
Alibaba's Qwen team pushed Qwen-Image-2.1 to Hugging Face and ModelScope today, pairing a 7B, 32-layer single-stream DiT with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE that outputs native 2048x2048 at 40 steps. The release ships two 9B PE-T2I/PE-I2I prompt-rewriter checkpoints, supports up to 10 reference images and mask/circle-based local edits, and switches licensing from Apache 2.0 on the earlier Qwen-Image line to a non-commercial Qwen Research License Agreement — commercial users now need a separate agreement.
StepFun launches 600B Step 5 Preview API at $1 in, $2.70 out
StepFun officially announced Step 5 Preview on September 20, a 600B-parameter sparse MoE with 27B active per token and a 1M-token context, and opened API access the same day. Artificial Analysis pegs the model at 44 on its Intelligence Index — matching Kimi K3 Max and roughly a seventh the price of GPT-5.6 Sol — at $1 per million input tokens and $2.70 per million output tokens with a 95% cache discount. StepFun says full open weights follow on October 15; the Hugging Face repo currently ships only a .gitattributes file.
UMG and Sony sue Suno again over 60,202 recordings in v6
Universal Music Group and Sony Music filed a 45-page complaint in the US District Court for Massachusetts on Friday accusing Suno's newly launched 'v6' family of models of being 'fruit of the same poisoned tree' — trained on outputs of Suno's earlier infringing models. The suit puts a number on the alleged violations: 60,202 copyrighted recordings, which the labels call 'only a small portion' of the total infringed. Warner, BMG, and Believe licensed content to power v6; UMG and Sony did not, and are seeking statutory damages and attorneys' fees.
Claude Opus 5 hacks OpenAI employee accounts hours after release
Three-person security startup Hacktron AI chained two OpenAI vulnerabilities — a memory bug in the libheif library reached through HEIF image uploads to the OpenAI community forum, which runs on Discourse — to take over multiple OpenAI employee ChatGPT and Codex accounts and open a pull request in OpenAI's internal monorepo. Claude Opus 4.8 had failed at the exploit across multiple sessions; within hours of Anthropic releasing Opus 5, the researchers succeeded. Less than 72 hours passed from initial discovery to internal repo access; OpenAI paid a $6,500 bug bounty and fixed the flaws.
Alibaba ships Qwen3.8-Omni-Flash omnimodal with 1M context
Alibaba's Qwen team released Qwen3.8-Omni-Flash on Sept 18, a native omnimodal model that jointly processes text, images, audio and video with a 1M-token context. On roughly 30 evaluations it beats Qwen3.5-Omni-Plus by 26%+ on average, with a 45.7% token-usage reduction on agentic video tasks. API-only via QwenCloud, Alibaba Cloud Model Studio and Qwen Studio at $0.15/M input and $0.47/M output; no open weights at launch.
Anthropic confirms Bay Area wet lab for AI-directed biology
Anthropic's head of life sciences Eric Kauderer-Abrams confirmed to Reuters that the company operates a Bay Area wet lab where Claude tests biological theories through physical experiments, marking a move beyond computer-based research. The lab focuses on fundamental biology rather than drug discovery, and this week Anthropic launched a Life Sciences Verification Program giving vetted researchers access to its most powerful models. The disclosure follows Anthropic's ~$400M April acquisition of stealth biotech Coefficient Bio.
Trump vows to form 'AI Force' and name an AI czar
In a Truth Social post Friday, President Trump said he will appoint an AI czar and stand up a new 'AI Force' modeled on the Space Force to oversee the industry, adding that 'only High I.Q. individuals need apply.' He dismissed AI safety concerns as a 'hoax' and vowed the White House 'will not in any way hinder or stifle' AI growth, framing US leadership as necessary to 'beat China.' No appointee was named.
PrismML's 5.9GB Ternary Bonsai 2 keeps 98.2% of Qwen3.8 27B
PrismML released Ternary Bonsai 2 27B, a ternary-quantized (−1/0/+1) rewrite of Alibaba's Qwen3.8 27B that shrinks the model from 54GB to 5.9GB — 9.1x compression — while retaining 98.2% of aggregate benchmark performance across 20 evals (99.5% math retention, 99.3% coding). The Apache 2.0 model uses blockwise Hadamard rotation at 1.71 bits/weight, runs at 142.5 tok/s on RTX 5090 and 46.8 tok/s on M5 Max, and fits on 16GB laptops via PrismML's llama.cpp fork.
finance.yahoo.com
2d ago
ALERTA 32
Anthropic revenue pace hits $100B, IPO delayed to November
NYT reports (via Axios/Yahoo Finance) that Anthropic is now pacing to generate over $100 billion in annualized revenue this year — up 50% from the $65 billion figure disclosed in July, and more than 10x end-of-2025 levels — driven by Claude Code and Cowork enterprise adoption. The company has pushed its IPO from October to November 2026 to include Q3 financials, targeting a ~$2 trillion valuation and potentially the largest IPO in history.