Amjad Masad当日精选
Amjad Masad 部署了一个基于微调LLM的国际象棋引擎,使用GRPO RL训练,目标达到2000+ Elo,目前估计1200 Elo,并提供了教程
AI 摘要
Amjad Masad 部署了一个基于微调LLM的国际象棋引擎,使用GRPO RL训练,目标达到2000+ Elo,目前估计1200 Elo,并提供了教程。
推荐理由常规快讯,保留列表
原文
1. One small finetuned LLM (no custom pretraining or architecture) 2. The model has to produce the moves with no chess engine assistance
If you relax these constraints, it gets much easier.
Amjad Masad: If you want to play my chess engine (WIP): https://qwen-chess.replit.app/
It already seems to perform better than frontier models on chess.
It’s fine-tuned on 2M stockfish-labeled positions, then short GRPO RL pass.
Documentation/tutorial with all the experiments and annotated code.
61/100
讨论
暂无评论。