Zhipu, a Chinese artificial intelligence company, has introduced a new AI model named GLM-5.3, which it claims surpasses models from Anthropic and OpenAI in its ability to identify software vulnerabilities. The company released benchmark data indicating that GLM-5.3 outperforms Fable 5 and GPT-5.6 Sol on the CyberGym benchmark, a test designed to evaluate an AI model's proficiency in resolving real-world cybersecurity challenges.
According to Zhipu, the model's "cyber capability developed faster than we expected" during post-training. The company stated that GLM-5.3 achieved state-of-the-art results on CyberGym for vulnerability discovery, showing significant improvements further along the exploitation chain. Zhipu emphasized that the model's advancements extended beyond identifying isolated flaws, enabling it to "reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains."
Zhipu reported collaborating with Chinese companies to test GLM-5.3 on actual codebases. These tests allegedly uncovered 2,436 vulnerabilities across 269 projects, with 1,097 of these issues categorized as medium-to-high severity. The identified vulnerabilities spanned various software components, including system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. The company noted that "Many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years."
While GLM-5.3 demonstrated strong performance in bug finding, Zhipu also acknowledged that the model performed less effectively than Western counterparts on other security and coding benchmarks. Despite this, the company highlighted the rapid development of GLM-5.3's bug-finding capabilities, suggesting that China is quickly closing any perceived gap in its ability to identify software weaknesses in rival systems, particularly following the debut of Anthropic’s Mythos.






