<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>强化学习 on zo0043</title><link>http://blog.zero43.top/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/</link><description>Recent content in 强化学习 on zo0043</description><image><title>zo0043</title><url>https://i.postimg.cc/7hwBy7VS/calcr.png</url><link>https://i.postimg.cc/7hwBy7VS/calcr.png</link></image><generator>Hugo -- 0.130.0</generator><language>zh</language><copyright>©2024 zo0043</copyright><lastBuildDate>Wed, 30 Sep 2026 16:40:00 +0800</lastBuildDate><atom:link href="http://blog.zero43.top/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/index.xml" rel="self" type="application/rss+xml"/><item><title>DeepSeek 论文课（八）：模型怎么会"想"了：GRPO 与 R1</title><link>http://blog.zero43.top/posts/deepseek-paper-course-08-rl-reasoning/</link><pubDate>Wed, 30 Sep 2026 16:40:00 +0800</pubDate><guid>http://blog.zero43.top/posts/deepseek-paper-course-08-rl-reasoning/</guid><description>模型怎么会&amp;quot;想&amp;quot;了：GRPO 与 R1 先给结论： 上一篇讲&amp;quot;结构省钱&amp;quot;，这一篇转向训练方法。模型的&amp;quot</description></item></channel></rss>