Zhipu AI launches and open-sources the "Niulai" model GLM-5.3-Flash, enhancing multimodal capabilities and inference cost efficiency.
Z.ai has released GLM-5.3-Flash, calling it the first native multimodal model in the GLM-5 series. The model features 320 billion total parameters and 18 billion active parameters, outperforming GLM-5.2 across multiple coding, agent, and vision benchmarks, and approaching Claude Opus 4.8 on certain coding tasks. GLM-5.3-Flash employs a hybrid architecture that combines sparse attention with linear attention, and introduces mechanisms such as mHC and IndexPool to reduce long-context inference costs and KV cache overhead.