Atomic Agent 在 GAIA Level 1 基准测试中以 69.8% 准确率击败 Hermes 的 58.5%,且速度快 1.6 倍,展示了开源 agent 运行时的性能优势
AI 摘要
Atomic Agent 是一款基于 MIT 许可证的开源 Agent 运行时。在 GAIA Level 1 基准测试中,使用相同的 4-bit qwen-3.6-35b 模型和 Apple M4 Max 硬件,Atomic Agent 以 69.8% 的准确率击败 Hermes 的 58.5%,且速度快 1.6 倍(3h12m vs 5h10m)。这一结果展示了开源 Agent 运行时的性能潜力,为开发者提供了更高效的选择。 核心观点: 1. Atomic Agent 在 GAIA Level 1 准确率 69.8%,比 Hermes 的 58.5% 高 11.3 个百分点。 2. 相同硬件和模型下,Atomic Agent 完成 53 任务耗时 3h12m,比 Hermes 快 1.6 倍。 3. Atomic Agent 解决 37/53 任务,Hermes 解决 31/53,多解决 6 个。
推荐理由高信息密度,值得细读
原文
Atomic Agent@atomicagent_ioJul 24
Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes
金句
Atomic Agent beat Hermes on GAIA: 69.8% vs. 58.5%, and finished 1.6× faster!
It’s an Open-Source-Agent-Runtime model under MIT license.
讨论
暂无评论。