<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>GRPO on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/grpo/</link>
    <description>Recent content in GRPO on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Mon, 20 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/grpo/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>论文精读｜ExploreVLA：密集世界建模与探索驱动的端到端自动驾驶</title>
      <link>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2604-02714/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2604-02714/</guid>
      <description>VLA 模型通过行为克隆学习驾驶策略，但受限于模仿学习无法探索专家分布之外的高质量策略。ExploreVLA 提出统一的理解-生成框架：用未来 RGB + 深度图生成作为密集世界建模目标，再利用世界模型的图像预测不确定性作为内在探索奖励，通过安全门控的 GRPO 优化策略。在 NAVSIM 上达到 93.7 PDMS 和 88.8 EPDMS。</description>
    </item>
    <item>
      <title>AutoVLA 代码讲解：把「开车」变成「说动作 token」并用 GRPO 微调</title>
      <link>https://auto-driving-blog.vercel.app/posts/code/autovla%E4%BB%A3%E7%A0%81%E8%AE%B2%E8%A7%A3/</link>
      <pubDate>Mon, 20 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/code/autovla%E4%BB%A3%E7%A0%81%E8%AE%B2%E8%A7%A3/</guid>
      <description>拆解 ucla-mobility/AutoVLA：动作 codebook 离散化、快慢思考两阶段训练、SFT &#43; GRPO 拒绝采样微调，看 VLA 如何端到端输出可执行驾驶动作</description>
    </item>
    <item>
      <title>论文精读：FeaXDrive — Feasibility-aware Trajectory-Centric Diffusion Planning</title>
      <link>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2604-12656/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2604-12656/</guid>
      <description>FeaXDrive 发现噪声中心扩散规划与轨迹可行性空间存在根本错位，提出轨迹中心扩散框架，通过自适应曲率正则化训练、可行驶区域引导推理和可行性感知 GRPO 后训练系统提升轨迹物理可行性。在 NAVSIM 上达 88.7 PDMS (IL) / 90.0 PDMS (RL)，曲率违反率从 8.59% 降至 0.88%。</description>
    </item>
    <item>
      <title>论文精读｜AutoDrive-P³：感知-预测-规划链式思维的统一强化微调——ICLR 2026 端到端驾驶新范式</title>
      <link>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2603-28116/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2603-28116/</guid>
      <description>当前 VLM 驾驶方案要么直接输出规划缺失 CoT 推理，要么将感知-预测-规划割裂为独立模块缺乏协同。AutoDrive-P³ 提出统一链式思维框架，通过 P³-CoT 数据集构建感知→预测→规划的结构化推理链，再用 P³-GRPO 算法进行分层渐进式强化微调——将奖励从规划反传到感知和预测模块，实现三模块联合优化。在 NAVSIM 上达到 89.9 EPDMS，nuScenes 上取得最低碰撞率。</description>
    </item>
  </channel>
</rss>
