Z.ai: GLM-5.3 Hits 84.5% on CyberGym Benchmark
Z.ai GLM-5.3 reaches 84.5% on CyberGym vulnerability benchmark LLM performance via agentic fine-tuning, surpassing GLM-5.2 and proprietary rivals while triggering safety delays.
SourceAnalysis
Z.ai pushed GLM-5.3 to 84.5% on the CyberGym vulnerability benchmark LLM performance test, topping leading proprietary models after a sharp lift from GLM-5.2. Engineers achieved the jump solely through targeted fine-tuning of the model’s agentic capabilities and optimization routines, leaving the base weights untouched. The upgraded system proved so effective at locating and exploiting vulnerabilities that Z.ai withheld open release pending further GLM model agentic capabilities fine-tuning safety reviews.
DeepLearning.AI
@DeepLearningAIWe are an education technology company with the mission to grow and connect the global AI community.