z.ai via Hacker News

Z.ai says GLM-5.3 built the inference stack that now serves it on 100,000+ Chinese chips

Summary

Z.ai published a technical account on Sept 17 saying an 'Infra Agent' powered by GLM-5.3 did most of the work to bring GLM-5.3-Flash (320B total / 18B active, 1M context) to production on a cluster of more than 100,000 Chinese-made AI accelerators — a scale it claims no one had operated on Chinese silicon before. The company reports end-to-end throughput tripled from baseline in under two weeks, with per-token cost and hardware efficiency 'comparable to mainstream Nvidia GPUs.' Z.ai frames the run as 'early forms' of recursive self-improvement, with caveats that humans still set objectives and boundaries.