智谱AI发布GLM-5.2,在移动应用开发能力上实现大幅提升,内部基准测试完成率从21/70提升至48/70
多源视角
RT Cunxiang Wang<br>GLM-5.2 is not only stronger on benchmarks, but also much better in real app development scenarios — iOS, Android, WeChat Mini Programs, and more.<br><br>Behind…
查看这条来源Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. <br><br>And more scenarios are on the way.<hr…
查看这条来源RT Zixuan Li<br>GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks.<br><br>Results:<br>- GLM-5.1: 21/70<br>- GLM…
查看这条来源RT OpenCode<br>GLM-5.2 now available in Go<br><br>text · 1M context · same pricing as 5.1
查看这条来源AI 摘要
智谱AI发布GLM-5.2模型,在移动应用开发能力上实现大幅提升。内部基准测试包含35个挑战性移动开发任务,每个任务运行两次共70次试验,以核心功能无重大问题为完成标准。GLM-5.2完成率从GLM-5.1的21/70提升至48/70,提升超过两倍,但仍低于Claude Fable 5的56/70。对开发者而言,该模型在长周期任务上的进步值得关注。 核心观点: 1. GLM-5.2在内部移动开发基准测试中完成率从21/70提升至48/70,提升超过两倍。 2. 基准测试包含35个挑战性移动开发任务,每个任务运行两次共70次试验。 3. Claude Fable 5在该基准上完成率为56/70,高于GLM-5.2。
推荐理由高信息密度,值得细读
原文
无法获取正文,可打开下方原始链接查看。
讨论
暂无评论。