NVIDIA AI当日精选
SGLang 正式支持 Thinky Machines 的 Inkling-Small 模型在双 DGX Spark 系统(通过 ConnectX-7 互联)上运行,实现本地强智能体能力,当前速度 24 tok/s(无 MTP),并提供了快速上手指南
AI 摘要
社区贡献者@realJerryzhou、@ispobaoke等实现了在紧凑平台上本地运行强智能体的支持,当前设置(并发1)推理速度24 tok/s(无MTP),并提供了cookbook快速上手指南。开发者可以基于此在本地部署强智能体,未来还将支持DSpark系统。 核心观点: 1. 当前推理速度24 tok/s(无MTP),在紧凑平台上本地运行强智能体。
推荐理由高信息密度,值得细读
原文
Huge thanks to @realJerryzhou @ispobaoke for implementing the support, with valuable community contributions from @u1tra_instinct and @TechMDAI.
This enables strong agentic capabilities to run locally on a compact platform. The current setup achieves 24 tok/s (concurrency=1) without MTP.
DSpark support is coming soon.
Check out the cookbook below to get started 👇
金句
This enables strong agentic capabilities to run locally on a compact platform.
The current setup achieves 24 tok/s (concurrency=1) without MTP.
Agent本地推理紧凑平台推理速度社区贡献DSpark70/100
讨论
暂无评论。