<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>🎯 DPO on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/-dpo/</link>
    <description>Recent content in 🎯 DPO on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Sun, 19 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/-dpo/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>知识点拆解｜RLHF 与偏好对齐详解：从大模型到自动驾驶驾驶风格学习</title>
      <link>https://auto-driving-blog.vercel.app/posts/knowledge/rlhf%E4%B8%8E%E5%81%8F%E5%A5%BD%E5%AF%B9%E9%BD%90%E8%AF%A6%E8%A7%A3/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/knowledge/rlhf%E4%B8%8E%E5%81%8F%E5%A5%BD%E5%AF%B9%E9%BD%90%E8%AF%A6%E8%A7%A3/</guid>
      <description>RLHF 通过人类偏好反馈将模型行为对齐到人类期望。本文详解 RLHF 三阶段流程、Bradley-Terry 奖励模型、DPO 直接偏好优化，并分析 GRPO 如何避免显式 reward 建模。在驾驶场景中，偏好对齐天然适配&amp;rsquo;不同人有不同驾驶风格&amp;rsquo;这一人类直觉——有人偏好激进、有人偏好保守，RLHF 让 VLA 模型学会按需生成驾驶行为。</description>
    </item>
  </channel>
</rss>
