Kimi K3在AA-Briefcase基准测试中取得第二高分(Elo 1543),但每任务成本高达10.57美元,是前代Kimi K2.6的10倍
AI 摘要
Kimi K3在AA-Briefcase基准测试中取得第二高分(Elo 1543),但每任务成本高达10.57美元,是前代Kimi K2.6的10倍。
推荐理由常规快讯,保留列表
原文
Last week @Kimi_Moonshot released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (1574)
AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables such as spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo based on correctness, analytical quality, and presentation quality
Key results for Kimi K3 on AA-Briefcase:
➤ Second only to Fable 5: Kimi K3 achieves an AA-Briefcase Elo of 1543, the second-highest score overall, ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). This is a +727 improvement over the previous-generation Kimi K2.6 (816) and puts Kimi K3 only behind Fable 5
➤ Strong objective and analytical performance, with comparatively weaker presentation: Kimi K3 achieves a rubric pass rate of 51%, second only to Claude Fable 5 (56%) and ahead of Claude Sonnet 5 (max, 42.3%) and GPT-5.6 Sol (max, 41.8%). It also records an analytical quality Elo of 1754, comparable to Claude Fable 5 (1744). Presentation quality is comparatively weaker, with a Presentation Elo of 1471, below GPT-5.6 Sol (max, 1660) and Claude Opus 4.8 (max, 1492)
➤ ~10x increase in Cost per Task: Kimi K3 averages a cost of $10.57 per task, a ~10x increase from Kimi K2.6, placing it among the most expensive models to run on AA-Briefcase. This is driven by model token pricing, increased output tokens and relatively high turn use, averaging 83 turns per task, versus 67 for Claude Fable 5 and 50 for GPT-5.6 Sol (max). Kimi K3 is priced at $3/$15 per 1M input/output tokens, with a 90% discount for cached tokens
➤ Averages nearly an hour per task: Kimi K3 has an average Time per AA-Briefcase Task of 56.4 minutes. This is driven by a high number of turns, high output token use, and lower speeds using the first-party Kimi API
讨论
暂无评论。