Z.ai Runs Its 320B-Parameter Model on 100,000+ Chinese Chips — and Tripled Throughput in Two Weeks
Z.ai reported that its GLM-5.3-Flash model — 320 billion parameters, 1 million token context — is running on more than 100,000 Chinese-made accelerators with throughput tripled in under two weeks.