# 作者指出旧基准已无法区分前沿模型，并介绍Karpathy用《指环王》文本生成Three.js世界的实验，推动基准进化

- 来源：Chubby
- 发布时间：2026-08-02 16:17
- AIWatch 分数：61
- AIWatch 标记：未精选
- AIWatch 链接：https://aiwatch.icu/events/evt_01kz0tchb49h7cdmc3nd1qt90p
- 原文链接：https://x.com/kimmonismus/status/2083829468032876616

## 精选理由

常规快讯，保留列表

## AI 摘要

作者指出旧基准已无法区分前沿模型，并介绍Karpathy用《指环王》文本生成Three.js世界的实验，推动基准进化。

## 正文

“Create an SVG of a pelican riding a bicycle” was once surprisingly helpful. Today, frontier models routinely produce convincing results, so the test tells us increasingly little about where the actual frontier is.

Karpathys experiment is a fascinating attempt to push the benchmark forward: give Opus 5 the opening of The Lord of the Rings, a massive token budget, and two hours to turn it into an interactive Three.js world. And voila!

This tests far more than one-shot generation: The model has to maintain coherence over thousands of lines of code, translate prose into a spatial system, coordinate objects and animations, and continuously inspect its own work.

Really love where the testing / benchmarks are moving to!

Love to see GPT-5.6 and the new DeepSeek flash on this one.
