<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>🏆 GAIL on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/-gail/</link>
    <description>Recent content in 🏆 GAIL on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Sun, 19 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/-gail/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>知识点拆解｜逆强化学习详解：从专家演示中学会&#39;为什么这么开&#39;</title>
      <link>https://auto-driving-blog.vercel.app/posts/knowledge/%E9%80%86%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0%E8%AF%A6%E8%A7%A3/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/knowledge/%E9%80%86%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0%E8%AF%A6%E8%A7%A3/</guid>
      <description>逆强化学习（IRL）解决的是&amp;rsquo;从行为反推动机&amp;rsquo;的问题——给定专家演示，推断 driving force 背后的 reward 函数。本文详解最大熵 IRL、GAIL/AIRL 对抗方法，阐明 IRL 与 BC、RL 的本质区别，并梳理 IRL 在自动驾驶中的应用：从人类驾驶数据中学习 reward，再用 RL 优化策略。最后讨论偏好 IRL 和 VLM 作为 reward 的最新趋势。</description>
    </item>
  </channel>
</rss>
