<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>多模态 on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.vercel.app/tags/%E5%A4%9A%E6%A8%A1%E6%80%81/</link>
    <description>Recent content in 多模态 on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.vercel.app/images/share.png</url>
      <link>https://auto-driving-blog.vercel.app/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Mon, 20 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.vercel.app/tags/%E5%A4%9A%E6%A8%A1%E6%80%81/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>论文精读：Drive-JEPA — Video JEPA 预训练 &#43; 多模态轨迹蒸馏的端到端驾驶</title>
      <link>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2601-22032/</link>
      <pubDate>Mon, 20 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.vercel.app/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2601-22032/</guid>
      <description>Drive-JEPA（2026）指出端到端驾驶的两个瓶颈：(1) 视频世界模型预训练收益有限；(2) 每场景只有一条人类轨迹，多模态监督稀缺。它用 V-JEPA（而非生成式世界模型）在 208 小时驾驶视频上自监督预训练 ViT 编码器，得到规划对齐的表征；再用基于仿真器的「多模态轨迹蒸馏」把多样伪教师轨迹蒸馏进 proposal-centric 规划器，并以动量感知选择抑制帧间抖动。NAVISIM v1 达 93.3 PDMS、v2 达 87.8 EPDMS 双榜 SOTA，仅前视相机+轻量 transformer 即在无感知设定下超此前 SOTA 3 PDMS。</description>
    </item>
  </channel>
</rss>
