# SGLang 正式支持 Thinky Machines 的 Inkling-Small 模型在双 DGX Spark 系统（通过 ConnectX-7 互联）上运行，实现本地强智能体能力，当前速度 24 tok/s（无 MTP），并提供了快速上手指南

- 来源：NVIDIA AI
- 发布时间：2026-08-01 08:30
- AIWatch 分数：70
- AIWatch 标记：当日精选
- AIWatch 链接：https://aiwatch.icu/events/evt_01kyxkvhwd6rqx7g5b1rrctda1
- 原文链接：https://x.com/NVIDIAAI/status/2083375013936476405

## 精选理由

高信息密度，值得细读

## AI 摘要

社区贡献者@realJerryzhou、@ispobaoke等实现了在紧凑平台上本地运行强智能体的支持，当前设置（并发1）推理速度24 tok/s（无MTP），并提供了cookbook快速上手指南。开发者可以基于此在本地部署强智能体，未来还将支持DSpark系统。
核心观点：
1. 当前推理速度24 tok/s（无MTP），在紧凑平台上本地运行强智能体。

## 正文

Huge thanks to @realJerryzhou @ispobaoke for implementing the support, with valuable community contributions from @u1tra_instinct and @TechMDAI.

This enables strong agentic capabilities to run locally on a compact platform. The current setup achieves 24 tok/s (concurrency=1) without MTP.

DSpark support is coming soon.

Check out the cookbook below to get started 👇
