Fireworks AI 与 MiniMax AI 公开了稀疏注意力内核仓库,通过优化加载和存储管道实现 1.6 倍吞吐量提升
多源视角
RT Fireworks<br>Weights aren't the only thing that should be open-source.<br><br>Both kernel repos behind our recent work with @MiniMax_AI are now public:<br>→ Fireworks: http://gi…
查看这条来源Weights aren't the only thing that should be open-source.<br><br>Both kernel repos behind our recent work with @MiniMax_AI are now public:<br>→ Fireworks: http://github.com/fw-ai/m…
查看这条来源AI 摘要
Fireworks AI 与 MiniMax AI 公开了稀疏注意力内核仓库,通过优化加载和存储管道实现 1.6 倍吞吐量提升。
推荐理由常规快讯,保留列表
原文
Both kernel repos behind our recent work with @MiniMax_AI are now public: → Fireworks: http://github.com/fw-ai/minimax-kernels → Minimax: http://github.com/MiniMax-AI/MSA/tree/fireworks-msa
RyanLee: Great to see @FireworksAI_HQ continuous optimizations on MiniMax Sparse Attention (MSA).
By refining the attention kernel’s load and store pipelines, they’ve achieved a 1.6x throughput uplift.
This work should bring meaningful value to open-source implementations.
讨论
暂无评论。