# Dylan 通过 8 种动物×6 种交通工具的 48 组提示词测试 7 个模型，未发现 AI 实验室存在“pelicanmaxxing”现象

- 来源：Simon Willison
- 发布时间：2026-07-23 07:01
- AIWatch 分数：58
- AIWatch 标记：未精选
- AIWatch 链接：https://aiwatch.icu/events/evt_01ky619dpwrat25e2sajh4vm1e
- 原文链接：https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything

## 精选理由

常规快讯，保留列表

## AI 摘要

Dylan 通过 8 种动物×6 种交通工具的 48 组提示词测试 7 个模型，未发现 AI 实验室存在“pelicanmaxxing”现象。

## 正文

Are AI labs pelicanmaxxing?

I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.

Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.

There's a neat filter view for exploring the results:

Screenshot of a grid for sample 1/3 of GLM-5.2, with pelicn and flamingo and heron riding bicycle, unicycle, skateboard, scooter, plane and boat

For the models he tested he could find no evidence of pelimaxxing:

The pelicans on bicycles don’t look any better Labs are not better at drawing pelicans Labs are not better at drawing bicycles Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty The pelican-bicycle scenes don’t look memorized [...]  Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.

Via Hacker News

Tags: ai, generative-ai, llms, evals, pelican-riding-a-bicycle
