第三方评测对比 Kimi K3 max 与 GPT-5.6 Sol max 在 DeepSWE 基准上的表现,显示两者各有所长,推荐级联使用以提升覆盖率和降低成本。
第三方评测在 DeepSWE 基准上对比 Kimi K3 max 与 GPT-5.6 Sol max,两者各有所长:Kimi 在多次尝试(pass@2/4)、成本、Rust 和运维领域占优;Sol 在单次(pass@1)、速度、Python/TS/JS 和序列化等领域占优。任务级相关低(0.46),失败模式不同,推荐级联使用。先 Kimi 再 Sol 的级联策略可达 85.6% 解题率,成本 $7.30/任务,比单独 Sol 覆盖高 13 个百分点且成本更低,为开发者提供高效低成本方案。
核心观点:
1. pass@2/4 上 Kimi 反超 GPT,多次尝试优势明显(Kimi pass@4 89.4% vs GPT 85.8%)。
2. Kimi 成本仅为 GPT 约 55%($4.65 vs $8.37),每 $100 解题数近 3 倍(14.7 vs 5.3)。
3. 级联策略(先 Kimi 再 Sol)实现 85.6% 解题率,成本 $7.30,比单独 Sol 覆盖高 13 个百分点。
xAI 发布 Grok Voice Think Fast 1.0 语音 API,在 EVA-Bench 评测中达到帕累托前沿,没有任何系统能在不牺牲体验时取得更高准确性,反之亦然。其具备类人时机、语调与温暖感,且定价仅为竞争对手的一小部分。开发者可在语音智能体场景中,以显著降低的成本获得无需权衡准确度与体验的商用语音交互能力。
核心观点:
1. Grok Voice Think Fast 1.0 在 EVA-Bench 达帕累托前沿,无系统能在不牺牲体验时击败其准确性,反之亦然。
2. Grok Voice Think Fast 1.0 定价为竞争对手一个 fraction,并提供类人时机、语调和温暖感。
同事件合集7 条来源
主线xAI
Grok Voice offers state-of-the-art performance with human-like timing, tone, and warmth. And it's a fraction the price of competitors. Check it out: h...
Grok Voice offer tate-of-the-art performance with human-like timing, tone, and warmth. And it' a fraction the price of competitor .<br><b
The Grok Build Plugin Marketplace is now in beta. Build with MongoDB, Vercel, Sentry, Cloudflare, and Chrome DevTools plugins from your terminal. Read...
The Grok Build Plugin Marketplace i now in beta. <br><br>Build with MongoDB, Vercel, Sentry, Cloudflare, and Chrome DevTool plugin from y
The Grok Build Plugin Marketplace is live. Build with MongoDB, Vercel, Sentry, Cloudflare, and Chrome DevTools plugins from your terminal. Read more h...
The Grok Build Plugin Marketplace i live. <br><br>Build with MongoDB, Vercel, Sentry, Cloudflare, and Chrome DevTool plugin from your ter
Read more about how Tori, eToro's agent, leverages models and real-time data from SpaceXAI to help consumers analyze market sentiment https://x.ai/new...
Read more about how Tori, eToro' agent, leverage model and real-time data from SpaceXAI to help con umer analyze market entiment<br><br