<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>🔗 多模态对齐 on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/-%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AF%B9%E9%BD%90/</link>
    <description>Recent content in 🔗 多模态对齐 on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Sun, 19 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/-%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AF%B9%E9%BD%90/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>知识点拆解｜多模态对齐与指令微调详解：从 VLM 到 VLA 的关键桥梁</title>
      <link>https://auto-driving-blog.vercel.app/posts/knowledge/%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AF%B9%E9%BD%90%E4%B8%8E%E6%8C%87%E4%BB%A4%E5%BE%AE%E8%B0%83%E8%AF%A6%E8%A7%A3/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/knowledge/%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AF%B9%E9%BD%90%E4%B8%8E%E6%8C%87%E4%BB%A4%E5%BE%AE%E8%B0%83%E8%AF%A6%E8%A7%A3/</guid>
      <description>多模态对齐是 VLA 的&amp;rsquo;关键桥梁&amp;rsquo;——它决定视觉信号能否被 LLM 正确理解。本文梳理从对比损失（CLIP）到 Q-Former（BLIP-2）到 MLP 投影（LLaVA）的对齐方案演进，以及 VLA 如何将对齐从&amp;rsquo;图像→文本&amp;rsquo;扩展到&amp;rsquo;图像→动作&amp;rsquo;，详解连续动作的 tokenization 策略。</description>
    </item>
  </channel>
</rss>
