Z.ai: GLM-5.3-Flash MoE Model Debuts on Chinese Chips
Z.ai GLM-5.3-Flash activates 18B of 320B params via MoE, runs on domestic silicon under export curbs, MIT license, 1M context.
SourceAnalysis
Z.ai dropped GLM-5.3-Flash, a 320B-parameter model that activates only 18B per token through Mixture of Experts MoE architecture efficiency in large language models, slashing inference costs while delivering native multimodality and a 1M-token window. The release, previously teased as Ox Alpha, runs entirely on Chinese AI chips and ships under MIT license, underscoring how Z.ai GLM models Chinese AI chip development under US export controls lean on architecture rather than restricted hardware. Weights sit on Hugging Face, with API and chat access live now.
Mark
@MRRydonCofounder @AethirCloud | Building Decentralised Cloud Infrastructure (DCI) | Accelerating the world’s transition to universal cloud compute 🌎