作者指出旧基准已无法区分前沿模型,并介绍Karpathy用《指环王》文本生成Three.js世界的实验,推动基准进化
AI 摘要
作者指出旧基准已无法区分前沿模型,并介绍Karpathy用《指环王》文本生成Three.js世界的实验,推动基准进化。
推荐理由常规快讯,保留列表
原文
“Create an SVG of a pelican riding a bicycle” was once surprisingly helpful. Today, frontier models routinely produce convincing results, so the test tells us increasingly little about where the actual frontier is.
Karpathys experiment is a fascinating attempt to push the benchmark forward: give Opus 5 the opening of The Lord of the Rings, a massive token budget, and two hours to turn it into an interactive Three.js world. And voila!
This tests far more than one-shot generation: The model has to maintain coherence over thousands of lines of code, translate prose into a spatial system, coordinate objects and animations, and continuously inspect its own work.
Really love where the testing / benchmarks are moving to!
Love to see GPT-5.6 and the new DeepSeek flash on this one.
讨论
暂无评论。