<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>🤔 个人思考 on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/-%E4%B8%AA%E4%BA%BA%E6%80%9D%E8%80%83/</link>
    <description>Recent content in 🤔 个人思考 on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Wed, 22 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/-%E4%B8%AA%E4%BA%BA%E6%80%9D%E8%80%83/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>个人思考｜强化学习 &#43; 生成式轨迹规划：四种范式、核心挑战与未来方向</title>
      <link>https://auto-driving-blog.vercel.app/posts/thoughts/rl-diffusion-traj/</link>
      <pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/thoughts/rl-diffusion-traj/</guid>
      <description>RL + 生成式轨迹是 2025-2026 最火热的方向之一。本文从一个小白的视角，梳理 SFT 的天花板、四种 RL 范式（奖励引导、拒绝采样、策略梯度、世界模型），深入分析 log-prob 难题和 reward hacking，用大量对比图和流程图让每个概念一目了然。最后给出了自己的理解和未来方向判断。</description>
    </item>
  </channel>
</rss>
