# 腾讯混元与人大高瓴学院开源 PlanningBench，提供 30 余项真实任务与自动验证，支持 LLM 规划能力系统评估与提升

- 来源：Tencent Hunyuan
- 发布时间：2026-06-05 15:46
- AIWatch 分数：71
- AIWatch 标记：当日精选
- AIWatch 链接：https://aiwatch.icu/events/evt_01ktvs832hfmf4af1zxhw264gf
- 原文链接：https://x.com/TencentHunyuan/status/2062803141314437391

## 精选理由

高信息密度，值得细读

## AI 摘要

大语言模型的规划能力是从“说”到“做”的关键。腾讯混元联合中国人民大学高瓴人工智能学院开源 PlanningBench 框架，提供 30 余项真实任务与自动验证能力，支持对 LLM 规划能力的可扩展评估与训练。该框架为开发者和研究者提供了系统化的基准测试环境，有助于更准确地衡量和提升大语言模型的实际规划与执行水平。
核心观点：
1. PlanningBench 涵盖 30 余项真实世界规划任务，用于评估大语言模型在实际场景中的规划表现。
2. 框架内置自动化验证机制，并同时支持对 LLM 规划能力的评估与训练。

## 正文

Tencent Hy, in collaboration with the Gaoling School of Artificial Intelligence at Renmin University of China, is excited to open-source PlanningBench - a scalable, verifiable framework for evaluating and training LLM planning capabilities.

With PlanningBench, you get:

✅ 30+ real-world planning tasks
✅ Automated verification
✅ Evaluation and training support

See how top-tier LLMs perform on PlanningBench 👇

Resources:
arXiv: https://arxiv.org/abs/2605.20873
GitHub: https://github.com/Tencent-Hunyuan/PlanningBench
HuggingFace: https://huggingface.co/datasets/tencent/PlanningBench

#PlanningBench #TencentHunyuan #OpenSource 📷
