🌐 Language: English · 简体中文
一个最前沿的视觉-语言模型论文、模型与代码仓库的综合整理与综述。
关于翻译范围:为了便于维护并保留技术术语的精确性,下方各表格中的论文标题、模型名称、数据集名称与链接均保持英文原文;前言、章节标题、表格列名、以及小节描述均译为中文。
视觉-语言模型(VLM)的设计在短短六年间经历了四个截然不同的架构时代 —— 而第三代又分化为两个并行的分支。早期模型保持冻结的视觉塔与语言塔,通过对比学习对齐(CLIP),或用一个可学习的连接器桥接到冻结的 LM(BLIP-2、Flamingo)。2023–2025 年的一代将预训练 LLM 作为主干,把视觉作为外挂适配器接入(LLaVA、Qwen2.5-VL、GPT-4V)。2025–2026 年的一代则完全抛弃了桥接器,将所有模态早期融合进一个统一的 Transformer —— 沿输出维度分叉为两支;而到了 2026 年,主干正在演变为能预测、能行动的世界模型:
- 第三代 a — 原生多模态输入 → 文本输出。 图像、视频、(有时)音频统一进入早期融合的 token 流,但生成仍然是自回归文本。这是当今通用旗舰模型采用的设计:Qwen3.5 / Qwen3.6、Gemma 4、Gemini 3、GPT-5.4、Phi-4-Reasoning-Vision、Claude Opus 4.6、Nemotron 3 Nano Omni。
- 第三代 b — 全模态统一 I/O。 同样的融合主干,再加上专用的图像 / 视频解码器(VAE / DiT / Flow-Matching)和/或音频编解码器解码头,使得模型也能生成图像、视频与语音 —— 通过自回归,或日益流行的离散扩散 / AR-扩散(LLaDA2.0-Uni、Mamoda2.5)。这是统一模型采用的设计:BAGEL、Qwen3.5-Omni、InternVL-U、Emu3 / Emu3.5、Erin 5.0、DeepSeek-Janus-Pro、LLaDA2.0-Uni、Mamoda2.5。纯生成的专用模型共享同一解码栈而不含理解侧 —— Sora 2、Veo 3、Kling 现已能生成带同步音频的视频,并且它们正是第四代世界模型的底座(DreamX-World 基于 Wan,OmniDreams 基于 Cosmos)。
- 第四代 — 世界-行动模型(2026 →)。 统一主干将行动作为一等模态纳入,并与环境闭环:模型预测未来观测、维持持久状态与空间记忆,并输出行动 —— 生成器、感知器与策略合而为一:Cosmos 3、Kairos、DreamX-World 1.0、OmniDreams。
图示阅读(从左至右)。 第一代采用双塔设计 —— 通过对比学习对齐(CLIP:没有生成式解码器),或经由可学习的跨模态桥(如 Q-Former)连接到冻结的 LM —— 仅文本输出。第二代以预训练 LLM 为中心;MLP/Resampler 将视觉 token 投影到 LLM 的词表空间,由 LLM 完成全部推理 —— 仍是文本输出。第三代 a 抛弃桥接器:图像、视频、音频与文本共享单一分词器/嵌入空间,并经过一个早期融合的 Transformer —— 但输出依然是自回归文本。第三代 b 保留这一融合主干,并加入解码头(图像/视频 DiT、VAE、音频编解码),使模型能原生输出文本、图像、视频和/或语音 —— 纯生成的视频模型(Sora 2、Veo 3、Kling)复用同一解码栈,如今还带有同步音频。第四代加入行动 token 流、持久状态与策略头,闭合"观测 → 行动 → 下一观测"的循环:世界-行动模型同时是生成器、感知器与策略。第三代 a、b 与第四代并存;选择基本上取决于"你需要模型生成多少 —— 以及是否要它行动?"。
下面我们汇编了精选的论文、模型和 GitHub 仓库,涵盖:
- 前沿视觉-语言模型 从最新到最早的 VLM 收录(持续更新新模型与基准)。
- 评估 VLM 评测基准及其对应工作的链接。
- 后训练 / 对齐 包括 RL、SFT 等 VLM 对齐方面的最新工作。
- 应用 VLM 在具身智能、机器人等领域的应用。
- 欢迎贡献相关的综述、观点与数据集。
VLM Trends 是本仓库的在线看板。本 README 记录已有的工作,而 VLM Trends 追踪今天有什么变化 —— 每日从 arXiv、Hugging Face、GitHub 与 Semantic Scholar 抓取新发布的模型、论文、基准与数据集,按公开的评分标准打分,并按主题与模型系列归类。它同时把本综述做成可检索的形式,并以图表呈现各研究方向随时间的变化趋势。
我们用带日期的小型综述追踪那些尚未折叠到主表中的新 VLM、基准与后训练方法:
📂 展开全部 9 期报告 — 最新:2026-08-10,评测重心从"能否看见"转向"能否行动与记忆"(29 条新条目)
-
📰
2026-08-10— 最新:评测不再只问模型能否看见 —— HumanCLAW(VLM 能否通过身体行动?)、GST-Bench(从视频建立全局空间意识)、ChronoVision(基于潜在状态重建的时序推理)、WorldExam(重"反应"而非"外观");"评判"本身成为研究问题:OSReward、ConfBench、TruthLens。模型侧:Qwen3.8-Max(2.4T · 95B 激活,Vision Arena 第 2)、DiffusionGemma(26B-A4B 扩散版 Gemma)、Hunyuan3D-Buffalo 1.0(统一 3D);以及 N₀-VTLA(触觉 VLA)、Metis、Ego2Robot、VideoCoCo、OmniPack —— 7 月 22 日以来 29 条新条目。 -
📰
2026-07-22— 世界模型成为评估器 —— GigaWorld-1 + WMBench、RoboWorld(与真实世界相关性 r = 0.989)、世界-行动模型教程;Gemma 4 技术报告(免编码器 12B)、PRA-GRPO(4B 模型 V-Star 93.2%)、VRRL(可训练的自我反思)、LingBot-VLA 2.0(60,000 小时语料)、ROSA(机器人工厂推理服务)、ISPA(KV 缓存削减 50%)、OmniFocus、MoHallBench / LongVQUBench / 科学可视化素养基准;以及 7 月前沿浪潮:Gemini 3.6 Flash、Kimi K3(2.8T 开源 MoE)、GPT-5.5 / GPT-5.6 Sol、Grok 4.5、Qwen3.7-Plus —— 6 月 27 日以来 19 条新条目。 -
📰
2026-06-27— 世界模型基础模型密集发布 —— Cosmos 3(NVIDIA 全模态家族:最佳开源 T2I/I2V + RoboArena 最佳策略)、Kairos(4B 边缘实时世界模型栈,胜过 14B)、DreamX-World 1.0(5B MIT 开源交互式世界模型);"持久状态核心"批判 + Echo-Memory;ZPPO(提示内教师胜过 GRPO)、Qwen-RobotManip(38,100 小时语料)、Supervise What Survives、VisCritic(GUI 视觉过程奖励)、HPP(长视频)、IMCBench(医疗对话安全)—— 6 月 23 日以来 11 条新条目。 -
📰
2026-06-23— 聚焦世界模型 —— NVIDIA OmniDreams(实时闭环驾驶世界模型)、Mirage(隐空间空间记忆)、Reward-as-Agent(面向世界模型的 GRPO)、WorldOlympiad 与 LongSpace-Bench(世界模型基准);以及 PP-OCRv6(34.5M 参数在 OCR 上超越 235B VLM)、Occ-VLM(3D 接地)、离散扩散 RL 推理、VLA 层剪枝、Hy-Embodied-0.5-VLA、RT-VLA(驾驶推理加速 44.8×)、VLA 语言引导 —— 6 月 2 日以来 12 条新条目。 -
📰
2026-06-02— Mamoda2.5(AR-扩散 DiT-MoE,编辑加速 95.9×)、VLM3(原生 3D 学习者)、AlphaGRPO(面向统一模型生成的 RL)、阶段式偏好优化、FastOCR / WindowQuant(KV 缓存高效化)、Fast-dDrive / CLOVER / CoWorld-VLA(驾驶 VLA)、Lost in Fog(推理一致性作为安全信号)、LiteGUI(免 SFT 的 GUI 智能体)、Health-Conditioned VLA、POLAR、TOC-Bench / VGenST-Bench(视频)、HalluCXR(医学)—— 5 月 16 日以来 16 条新条目。 -
📰
2026-05-16— LensVLM(Apple)、Nemotron 3 Nano Omni(NVIDIA)、LLaDA2.0-Uni、PLaMo 2.1-VL、S-GRPO / Faithful GRPO / GRPO-TTA / OpenSearch-VL、MindVLA-U1(超越人类驾驶)、VLADriver-RAG、Green-VLA、Anticipation-VLA、VLA Foundry、LAMO、ScreenExplorer、VideoZeroBench、Video-Oasis、MedThinkVQA、数据筛选实现 87× 更低算力 —— 4 月 28 日以来 34 条新条目。 -
📰
2026-04-28— Qwen3.6-27B & Qwen3.6-35B-A3B、Claude Mythos(受限预览)、S1-VL、GLM-5V-Turbo、FreshPER / GMPO / ARPO / GRPO-VPS、QUOTA、Fast-dVLM、VLA-World、SpanVLA、VLA-Forget、R-VLM、UILoop、WebForge、WorldMark、Video-MME-v2、CrossMath、BabyVision、SlowBA —— 4 月 13 日以来 30 条新条目。 -
📰
2026-04-13— LFM2.5-VL-450M、EXAONE 4.5、Gemma 4、Granite 4.0 3B Vision、InternVL-U、GLM-4.6V、Vero、MolmoWeb、UniDriveVLA、QAPruner、Firebolt-VL、CoME-VL 等。 -
📰
2026-03-25— GPT-5.4、Phi-4-Reasoning-Vision-15B、Gemini 3.0、Qwen3.5、Claude Opus 4.6、Molmo2 等。
欢迎贡献与讨论!
🤩 标有 ⭐️ 的论文由本仓库的维护者贡献。如果您觉得有用,欢迎给本仓库 Star 或引用我们的论文。
-
- 1.1. 🌍 世界模型
-
- 2.1. 大规模预训练与后训练数据集
- 2.2. VLM 数据集与评估
- 2.3. 具身 VLM 的基准、仿真器与生成模型
-
-
🔥 后训练 / 对齐 / 提示工程 🔥
- 3.1. VLM 强化学习对齐
- 3.2. 常规微调 (SFT)
- 3.3. VLM 对齐相关 GitHub 仓库
- 3.4. 提示工程
-
@InProceedings{Li_2025_CVPR,
author = {Li, Zongxia and Wu, Xiyang and Du, Hongyang and Liu, Fuxiao and Nghiem, Huy and Shi, Guangyao},
title = {A Survey of State of the Art Large Vision Language Models: Benchmark Evaluations and Challenges},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops},
month = {June},
year = {2025},
pages = {1587-1606}
}
表格按发布时间从新到旧排序。各列依次为:模型 · 年份 · 架构 · 训练数据 · 参数量 · 视觉编码器/分词器 · 预训练主干。
| 模型 | 年份 | 架构 | 训练数据 | 参数量 | 视觉编码器/分词器 | 预训练主干 |
|---|---|---|---|---|---|---|
| Qwen3.8-Max (Alibaba) | 08/03/2026 | Sparse MoE + hybrid attention; text + vision in, 1M context; Vision Arena #2 | Undisclosed | 2.4T total · 95B active | Native multimodal | Qwen3.8 |
| DiffusionGemma (Google) | 08/05/2026 | Diffusion (non-autoregressive) language model in the Gemma family | Undisclosed | 26B total · 4B active | Native multimodal | Gemma |
| Hunyuan3D-Buffalo 1.0 (Tencent) | 08/05/2026 | Unified multimodal — 3D generation + understanding + editing | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Gemini 3.6 Flash (Google) | 07/21/2026 | Decoder-only / natively multimodal input (text, image, speech, video → text) | Undisclosed | Undisclosed | Native multimodal | Gemini 3.x |
| Kimi K3 (Moonshot AI) | 07/17/2026 | MoE, natively multimodal reasoning — text + image + video in one trunk (no separate vision module); 1M context; open weights due 07/27/2026 | Undisclosed | ~2.8T total (MoE) | Native (encoder-integrated) | New architecture |
| GPT-5.6 Sol (OpenAI) | 07/09/2026 | Decoder-only; text + image in, 1.05M context / 128K output; max & ultra reasoning modes, sub-agent orchestration (Ultra); family: Luna / Terra / Sol | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Grok 4.5 (xAI) | 07/08/2026 | Decoder-only reasoning (extended thinking); text + image + files in, 500K context | Undisclosed | ~1.5T (reported) | Undisclosed | Undisclosed |
| Qwen3.7-Plus (Alibaba) | 06/01/2026 | Natively multimodal agent — image + video understanding, GUI grounding, tool invocation (note: Qwen3.7-Max is text-only) | Undisclosed | Undisclosed | Native multimodal ViT | Qwen3.7 |
| Mamoda2.5 (InclusionAI) | 05/04/2026 | AR-Diffusion + DiT-MoE (unified understanding + generation; 128 experts, Top-8) | Multimodal und. + image/video generation & editing | 25B total · 3B active | Semantic tokenizer + DiT decoder head | Mamoda2 |
| Nemotron 3 Nano Omni (NVIDIA) | 04/28/2026 | Hybrid MoE (omni-modal: vision + audio + text) | Vision + audio + text joint training | 30B total · 3B active | Dynamic-res ViT + Conv3D temporal | Nemotron 3 |
| GPT-5.5 (OpenAI) | 04/23/2026 | Decoder-only; text + image in (GPT-5 input stack), computer-use screen reading in Codex; topped AA Intelligence Index at release | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Qwen3.6-27B (Alibaba) | 04/22/2026 | Decoder-only / natively multimodal input (thinking + non-thinking) | Multimodal pretraining + agentic mid-training | 27B dense | Native multimodal ViT | Qwen3.6 |
| Qwen3.6-35B-A3B (Alibaba) | 04/15/2026 | MoE / natively multimodal input | Multimodal pretraining + agentic SFT/RL | 35B total · 3B active | Native multimodal ViT | Qwen3.6 |
| LFM2.5-VL-450M (Liquid AI) | 04/11/2026 | Liquid Foundation Model | Undisclosed | 450M | Non-overlapping tile ViT | LFM2.5 |
| EXAONE 4.5 (LG AI Research) | 04/09/2026 | Unified VL | Undisclosed | 33B | Proprietary vision encoder | EXAONE 4.5 |
| Claude Mythos (Anthropic, gated preview) | 04/07/2026 | Decoder-only (frontier; Project Glasswing gated) | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Gemma 4 (Google) | 04/02/2026 | Decoder-only / MoE | Undisclosed (140+ languages) | E2B / E4B / 26B MoE / 31B Dense | Native multimodal | Gemini 3 |
| Granite 4.0 3B Vision (IBM) | 04/01/2026 | Decoder-only | Enterprise document corpora | 3B | Undisclosed | Granite 4.0 |
| GLM-5V-Turbo (Zhipu / Z.AI) | 04/01/2026 | Natively multimodal (vision-coding) with Multi-Token Prediction | 30+ task joint RL | Undisclosed | CogViT | GLM-5 |
| InternVL-U (Shanghai AI Lab) | 03/10/2026 | Unified (MLLM + MMDiT) | Multimodal understanding + generation | 4B | InternViT | InternVL |
| GPT-5.4 / GPT-5.4 Thinking (OpenAI) | 03/06/2026 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Phi-4-Reasoning-Vision-15B (Microsoft) | 03/04/2026 | Decoder-only | Curated synthetic + filtered data | 15B | High-res dynamic-resolution ViT | Phi-4 |
| Gemini 3.0 (Google) | 03/2026 | Unified Model | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Qwen3.5 (Alibaba) | 02/16/2026 | Unified VL (early fusion) | Trillions of multimodal tokens | 0.8B–397B (MoE, 17B active) | ViT (native) | Qwen3.5 |
| Claude Opus 4.6 (Anthropic) | 02/2026 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Erin 5.0 (Baidu) | 02/05/2026 | Unified Model (Visual, Text, Audio) | Unified Modality Dataset | - | CNN–ViT (Understanding)/Next-Frame-and-Scale Prediction (Generation) | Unified Autoregressive Transformer |
| Molmo2 (Allen AI) | 01/15/2026 | Decoder-only | 7 new video + 2 multi-image datasets (9.19M videos) | 4B / 7B / 8B | Bi-directional attention ViT | Qwen 3 / OLMo |
| Gemini 3 | 11/18/2025 | Unified Model | Undisclosed | - | - | - |
| Emu3.5 | 10/30/2025 | Deconder-only | Unified Modality Dataset | - | SigLIP | Qwen3 |
| DeepSeek-OCR | 10/20/2025 | Encoder-Deconder | 70% OCR, 20% general vision, 10% text-only | 3B | DeepEncoder | DeepSeek-3B |
| Qwen3-VL | 10/11/2025 | Decoder-Only | - | 8B/4B | ViT | Qwen3 |
| Qwen3-VL-MoE | 09/25/2025 | Decoder-Only | - | 235B-A22B | ViT | Qwen3 |
| Qwen3-Omni (Visual/Audio/Text) | 09/21/2025 | - | Video/Audio/Image | 30B | ViT | Qwen3-Omni-MoE-Thinker |
| LLaVA-Onevision-1.5 | 09/15/2025 | - | Mid-Training-85M & SFT | 8B | Qwen2VLImageProcessor | Qwen3 |
| InternVL3.5 | 08/25/2025 | Decoder-Only | multimodal & text-only | 30B/38B/241B | InternViT-300M/6B | Qwen3 / GPT-OSS |
| SkyWork-Unipic-1.5B | 07/29/2025 | - | image/video.. | - | - | - |
| Grok 4 | 07/09/2025 | - | image/video.. | 1-2 Trillion | - | - |
| Kwai Keye-VL (Kuaishou) | 07/02/2025 | Decdoer-only | image/video.. | 8B | ViT | QWen-3-8B |
| OmniGen2 | 06/23/2025 | Decdoer-only & VAE | LLaVA-OneVision/ SAM-LLaVA.. | - | ViT | QWen-2.5-VL |
| Gemini-2.5-Pro | 06/17/2025 | - | - | - | - | - |
| GPT-o3/o4-mini | 06/10/2025 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Mimo-VL (Xiaomi) | 06/04/2025 | Decdoer-only | 24 Trillion MLLM tokens | 7B | [Qwen2.5-ViT | Mimo-7B-base |
| BAGEL (Bytedance) | 05/20/2025 | Unified Model | Video/Image/Text | 7B | SigLIP2-so400m/14](https://arxiv.org/abs/2502.14786) | Qwen2.5 |
| BLIP3-o | 05/14/2025 | Decdoer-only | (BLIP3-o 60K) GPT-4o Generated Image Generation Data | 4/8B | ViT | QWen-2.5-VL |
| InternVL-3 | 04/14/2025 | Decdoer-only | 200 Billion Tokens | 1/2/8/9/14/38/78B | ViT-300M/6B | InterLM2.5/QWen2.5 |
| LLaMA4-Scout/Maverick | 04/04/2025 | Decdoer-only | 40/20 Trillion Tokens | 17B | MetaClip | LLaMA4 |
| Qwen2.5-Omni | 03/26/2025 | Decdoer-only | Video/Audio/Image/Text | 7B | Qwen2-Audio/Qwen2.5-VL ViT | End-to-End Mini-Omni |
| QWen2.5-VL | 01/28/2025 | Decdoer-only | Image caption, VQA, grounding agent, long video | 3B/7B/72B | Redesigned ViT | Qwen2.5 |
| GLM-4.6V (Zhipu / Z.AI) | 12/2025 | Decoder-only | Undisclosed | 106B / 9B (Flash) | Undisclosed | GLM-4.6 |
| Ola | 2025 | Decoder-only | Image/Video/Audio/Text | 7B | OryxViT | Qwen-2.5-7B, SigLIP-400M, Whisper-V3-Large, BEATs-AS2M(cpt2) |
| Ocean-OCR | 2025 | Decdoer-only | Pure Text, Caption, Interleaved, OCR | 3B | NaViT | Pretrained from scratch |
| SmolVLM | 2025 | Decoder-only | SmolVLM-Instruct | 250M & 500M | SigLIP | SmolLM |
| DeepSeek-Janus-Pro | 2025 | Decoder-only | Undisclosed | 7B | SigLIP | DeepSeek-Janus-Pro |
| Inst-IT | 2024 | Decoder-only | Inst-IT Dataset, LLaVA-NeXT-Data | 7B | CLIP/Vicuna, SigLIP/Qwen2 | LLaVA-NeXT |
| DeepSeek-VL2 | 2024 | Decoder-only | WiT, WikiHow | 4.5B x 74 | SigLIP/SAMB | DeepSeekMoE |
| xGen-MM (BLIP-3) | 2024 | Decoder-only | MINT-1T, OBELICS, Caption | 4B | ViT + Perceiver Resampler | Phi-3-mini |
| TransFusion | 2024 | Encoder-decoder | Undisclosed | 7B | VAE Encoder | Pretrained from scratch on transformer architecture |
| Baichuan Ocean Mini | 2024 | Decoder-only | Image/Video/Audio/Text | 7B | CLIP ViT-L/14 | Baichuan |
| LLaMA 3.2-vision | 2024 | Decoder-only | Undisclosed | 11B-90B | CLIP | LLaMA-3.1 |
| Pixtral | 2024 | Decoder-only | Undisclosed | 12B | CLIP ViT-L/14 | Mistral Large 2 |
| Qwen2-VL | 2024 | Decoder-only | Undisclosed | 7B-14B | EVA-CLIP ViT-L | Qwen-2 |
| NVLM | 2024 | Encoder-decoder | LAION-115M | 8B-24B | Custom ViT | Qwen-2-Instruct |
| Emu3 | 2024 | Decoder-only | Aquila | 7B | MoVQGAN | LLaMA-2 |
| Claude 3 | 2024 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| InternVL | 2023 | Encoder-decoder | LAION-en, LAION- multi | 7B/20B | Eva CLIP ViT-g | QLLaMA |
| InstructBLIP | 2023 | Encoder-decoder | CoCo, VQAv2 | 13B | ViT | Flan-T5, Vicuna |
| CogVLM | 2023 | Encoder-decoder | LAION-2B ,COYO-700M | 18B | CLIP ViT-L/14 | Vicuna |
| PaLM-E | 2023 | Decoder-only | All robots, WebLI | 562B | ViT | PaLM |
| LLaVA-1.5 | 2023 | Decoder-only | COCO | 13B | CLIP ViT-L/14 | Vicuna |
| Gemini | 2023 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| GPT-4V | 2023 | Decoder-only | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| BLIP-2 | 2023 | Encoder-decoder | COCO, Visual Genome | 7B-13B | ViT-g | Open Pretrained Transformer (OPT) |
| Flamingo | 2022 | Decoder-only | M3W, ALIGN | 80B | Custom | Chinchilla |
| BLIP | 2022 | Encoder-decoder | COCO, Visual Genome | 223M-400M | ViT-B/L/g | Pretrained from scratch |
| CLIP | 2021 | Encoder-decoder | 400M image-text pairs | 63M-355M | ViT/ResNet | Pretrained from scratch |
2026 年中,世界模型从研究演示走向可发布的工程产物:统一主干同时充当生成器、感知器与策略。下表先列基础模型发布,随后是记忆 / 状态分析与基准。
| 模型 | 日期 | 类型 | 规模 / 许可 | 关键结果 | 链接 |
|---|---|---|---|---|---|
| HelloWorld | 08/06/2026 | Video world model with socially interactive characters | — | Pushes world models past physics and navigation into social dynamics: characters that respond to the viewer rather than merely persisting | Paper |
| VideoCoCo | 07/31/2026 | Agentic dual-engine text-to-video with Code-as-CoT | — | Infers temporal evolution symbolically via generated code instead of implicitly, then renders — a hybrid attack on physics violations in T2V | Paper |
| StatePlay | 07/29/2026 | State-aware game world model | — | Enforces mechanics consistency through explicit state, an answer to the persistent-state critique below | Paper |
| Cosmos 3 (NVIDIA) | 06/01/2026 | Omnimodal world-model family (language, image, video, audio, action) — mixture-of-transformers | Family; open code/checkpoints/data (OpenMDW-1.1) | SoTA across VL, video generation, robot policy; best open T2I/I2V + best RoboArena policy | Paper |
| Kairos | 06/15/2026 | Native world-model stack (understanding + generation + prediction) | 4B unified; Hybrid Linear Temporal Attention | Cross-embodiment data curriculum; real-time rollout on edge hardware; beats 14B on embodied benchmarks | Paper |
| DreamX-World 1.0 (Alibaba AMAP) | 06/15/2026 | General-purpose interactive T/I-to-video world model | 5B (Wan2.2-T2V-5B base); MIT open weights | Camera navigation, revisit consistency, promptable events; E-PRoPE encoding; few-step AR via causal forcing + DMD distillation | HF · Code |
| NVIDIA OmniDreams | 06/02/2026 | Real-time generative world model for closed-loop AV simulation | Mid/post-trained from Cosmos on 21k h driving | Action-conditioned sensor video; derived world-action model beats VLA-based Alpamayo 1.5 at 1/5 params | Paper |
| Mirage | 06/08/2026 | Latent spatial memory for video world models | — | 3D scene memory in diffusion latent space via depth-guided back-projection; faster + lighter than explicit-3D baselines | Paper |
| Reward as An Agent | 06/18/2026 | RL post-training for embodied world models | — | DynDiff-GRPO diversifies action-space exploration; agentic reward verification curbs reward hacking | Paper |
基准与分析
| 标题 | 日期 | 结论 | 链接 |
|---|---|---|---|
| WorldExam | 08/04/2026 | Separates apparent appearance from inherent reactivity: a world model can look right while reacting wrong, formalising the critique that visual realism has been over-weighted | Paper |
| WorldOlympiad | 06/09/2026 | Physics / geometry / interaction "triathlon": SoTA world models show substantial gaps in physical reasoning, 3D consistency, long-horizon control | Paper |
| Current World Models Lack a Persistent State Core | 06/18/2026 | World models treat the world as a "tracking shot" — off-screen entities freeze instead of evolving; persists across architectures and scales | Paper |
| Echo-Memory | 06/08/2026 | Controlled memory study: raw context beats compressed memory for capacity; state-space recurrence best for revisit consistency | Paper |
| LongSpace / LongSpace-Bench | 06/04/2026 | Video MLLMs fail long-horizon spatial recall without explicit spatial memory (3D cues + layer-aware retrieval) | Paper |
另见 §2.3 中更早的世界模型条目(UniSim、GAIA-1、LWM、Genesis、RoboGen),以及
2026-06-23/2026-06-27两期渐进研究报告。
用于 VLM 预训练 / 后训练的大规模多模态语料。
| 数据集 | 任务 | 规模 |
|---|---|---|
| MolmoWebMix (Allen AI)(04/2026) | Web Agent Training Trajectories | 100K+ synthetic + 30K human demos |
| Vero-600K(04/2026) | Broad Visual Reasoning RL Training | 600K samples from 59 datasets, 6 task categories |
| BigEarthNet.txt(03/2026) | Multi-sensor Earth Observation Image-Text | 464K images, 9.6M text annotations |
| OmniScience(02/2026) | Scientific Image Understanding | 1.5M figure-caption-context triplets |
| MaD-Mix(02/2026) | Multi-modal Data Mixture Optimization | Framework (0.5B–7B scale) |
| OVID(2026) | Open Video Pre-training | 10M hours, 300M frame-caption pairs |
| Molmo2 Video Datasets(01/2026) | Video Captions, QA, Tracking, Pointing | 9.19M videos (7 video + 2 multi-image datasets) |
| MMFineReason(/1/30/2026) | REasoning | 1.8M |
| FineVision(09/04/2025) | Mixed Domain | 24.3 M/4.48TB |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| MathVision | Visual Math | MC / Answer Match | Human | 3.04 | Repo |
| MathVista | Visual Math | MC / Answer Match | Human | 6 | Repo |
| MathVerse | Visual Math | MC | Human | 4.6 | Repo |
| VisNumBench | Visual Number Reasoning | MC | Python Program generated/Web Collection/Real life photos | 1.91 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| ROVER | Reciprocal Cross-Modal Reasoning | Visual Gen + Verbal Gen Eval | Human | 1.3 (1,876 images) | Paper |
| RealUnify | Math, World knowledge, Image Gen | Direct & StepWise Eval (Sec 3.3) | Script & Humanverification | 1.0 | |
| Uni-MMMU | Science, Code, Image Gen | DreamSim (Image Gen Eval) & String Matching (Understanding Eval) | - | 1.0 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| GST-Bench | Global Spatial Awareness from Continuous Video | Long-traversal scene integration (vs. single-viewpoint local perception) | - | - | Paper |
| ChronoVision | Multi-step Temporal Reasoning | Latent state reconstruction | - | - | Paper |
| MMOU | Omni-modal Long Video Understanding | MC | Human | 15 (9,038 videos) | Paper |
| Video-MMMU | Knowledge Acquisition from Professional Videos | MC + Knowledge Gain | Expert | 0.9 (300 videos) | Paper |
| MMVU | Expert-Level Multi-Discipline Video Understanding | MC | Expert | 3 (27 subjects) | Paper |
| VideoHallu | Video Understanding | LLM Eval | Human | 3.2 | |
| Video SimpleQA | Video Understanding | LLM Eval | Human | 2.03 | Repo |
| MovieChat | Video Understanding | LLM Eval | Human | 1 | Repo |
| Perception‑Test | Video Understanding | MC | Crowd | 11.6 | Repo |
| VideoMME | Video Understanding | MC | Experts | 2.7 | Site |
| EgoSchem | Video Understanding | MC | Synth / Human | 5 | Site |
| Inst‑IT‑Bench | Fine‑grained Image & Video | MC & LLM | Human / Synth | 2 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| VisionArena | Multimodal Conversation | Pairwise Pref | Human | 23 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| HumanCLAW | Can a VLM Act Through a Body? | Decouples the decision from motor control so failures attribute to perception, planning, or actuation | - | - | Paper |
| PerceptionBench | Atomic Visual Perception | Decomposes perception into primitive operations rather than end-task accuracy | - | - | Paper |
| C$^3$PO | Cross-Modal Composition & Counterfactuals (omni-modal) | Compositional and counterfactual probes for any-to-any models | - | - | Paper |
| OmniEarth | Geospatial / Remote Sensing VLM Eval | MC + Open VQA | Human (verified) | 44.2 (9,275 images, 28 tasks) | Paper |
| MultiHaystack | Multimodal Retrieval & Reasoning | Retrieval + QA | Human | 0.75 (46K+ candidates) | |
| DatBench | Discriminative, Faithful VLM Eval | MC (format-aware) | Synth | - | |
| MMLU | General MM | MC | Human | 15.9 | |
| MMStar | General MM | MC | Human | 1.5 | Site |
| NaturalBench | General MM | Yes/No, MC | Human | 10 | HF |
| PHYSBENCH | Visual Math Reasoning | MC | Grad STEM | 0.10 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| EMMA | Visual Reasoning | MC | Human + Synth | 2.8 | Repo |
| MMTBENCH | Visual Reasoning & QA | MC | AI Experts | 30.1 | Repo |
| MM‑Vet | OCR / Visual Reasoning | LLM Eval | Human | 0.2 | Repo |
| MM‑En/CN | Multilingual MM Understanding | MC | Human | 3.2 | Repo |
| GQA | Visual Reasoning & QA | Answer Match | Seed + Synth | 22 | Site |
| VCR | Visual Reasoning & QA | MC | MTurks | 290 | Site |
| VQAv2 | Visual Reasoning & QA | Yes/No, Ans Match | MTurks | 1100 | Repo |
| MMMU | Visual Reasoning & QA | Ans Match, MC | College | 11.5 | Site |
| MMMU-Pro | Visual Reasoning & QA | Ans Match, MC | College | 5.19 | Site |
| R1‑Onevision | Visual Reasoning & QA | MC | Human | 155 | Repo |
| VLM²‑Bench | Visual Reasoning & QA | Ans Match, MC | Human | 3 | Site |
| VisualWebInstruct | Visual Reasoning & QA | LLM Eval | Web | 0.9 | Site |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| ExtractBench | Schema-Guided Enterprise Document Extraction | Scored against a target schema rather than free-form answers | - | - | Paper |
| ConfBench | Confidence Calibration on Document Extraction | Asks whether the model knows when it is wrong, not only whether it is right | - | - | Paper |
| TableVision | Spatially Grounded Table Reasoning | 3-level Cognitive Eval | Human | 6.8 (13 sub-categories) | Paper |
| TextVQA | Visual Text Understanding | Ans Match | Expert | 28.6 | Repo |
| DocVQA | Document VQA | Ans Match | Crowd | 50 | Site |
| ChartQA | Chart Graphic Understanding | Ans Match | Crowd / Synth | 32.7 | Repo |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| FilmBench | Film-Grade Cinematic Video Generation | Craft criteria — shot grammar, continuity, staging | - | - | Paper |
| MPIE-Bench | Anatomically Plausible Multi-Person Interaction Editing | Anatomical plausibility under multi-subject edits | - | - | Paper |
| MSCOCO‑30K | Text‑to‑Image | BLEU, ROUGE, Sim | MTurks | 30 | Site |
| GenAI‑Bench | Text‑to‑Image | Human Rating | Human | 80 | HF |
| 数据集 | 任务 | 评估协议 | 标注者 | 规模 (K) | 代码 / 站点 |
|---|---|---|---|---|---|
| HallusionBench | Hallucination | Yes/No | Human | 1.13 | Repo |
| POPE | Hallucination | Yes/No | Human | 9 | Repo |
| CHAIR | Hallucination | Yes/No | Human | 124 | Repo |
| MHalDetect | Hallucination | Ans Match | Human | 4 | Repo |
| Hallu‑Pi | Hallucination | Ans Match | Human | 1.26 | Repo |
| HallE‑Control | Hallucination | Yes/No | Human | 108 | Repo |
| AutoHallusion | Hallucination | Ans Match | Synth | 3.129 | Repo |
| BEAF | Hallucination | Yes/No | Human | 26 | Site |
| GAIVE | Hallucination | Ans Match | Synth | 320 | Repo |
| HalEval | Hallucination | Yes/No | Crowd / Synth | 2 | Repo |
| AMBER | Hallucination | Ans Match | Human | 15.22 | Repo |
涵盖机器人导航 / 操作仿真器、自动驾驶、世界模型等。
| 基准 | 领域 | 类型 | 项目 |
|---|---|---|---|
| OSReward | Computer-Use Agents | Cross-platform reward-model evaluation | Paper |
| Drive-Bench | Embodied AI | Autonomous Driving | Website |
| Habitat, Habitat 2.0, Habitat 3.0 | Robotics (Navigation) | Simulator + Dataset | Website |
| Gibson | Robotics (Navigation) | Simulator + Dataset | Website, Github Repo |
| iGibson1.0, iGibson2.0 | Robotics (Navigation) | Simulator + Dataset | Website, Document |
| Isaac Gym | Robotics (Navigation) | Simulator | Website, Github Repo |
| Isaac Lab | Robotics (Navigation) | Simulator | Website, Github Repo |
| AI2THOR | Robotics (Navigation) | Simulator | Website, Github Repo |
| ProcTHOR | Robotics (Navigation) | Simulator + Dataset | Website, Github Repo |
| VirtualHome | Robotics (Navigation) | Simulator | Website, Github Repo |
| ThreeDWorld | Robotics (Navigation) | Simulator | Website, Github Repo |
| VIMA-Bench | Robotics (Manipulation) | Simulator | Website, Github Repo |
| VLMbench | Robotics (Manipulation) | Simulator | Github Repo |
| CALVIN | Robotics (Manipulation) | Simulator | Website, Github Repo |
| GemBench | Robotics (Manipulation) | Simulator | Website, Github Repo |
| WebArena | Web Agent | Simulator | Website, Github Repo |
| UniSim | Robotics (Manipulation) | Generative Model, World Model | Website |
| GAIA-1 | Robotics (Automonous Driving) | Generative Model, World Model | Website |
| LWM | Embodied AI | Generative Model, World Model | Website, Github Repo |
| Genesis | Embodied AI | Generative Model, World Model | Github Repo |
| EMMOE | Embodied AI | Generative Model, World Model | Paper |
| RoboGen | Embodied AI | Generative Model, World Model | Website |
| UnrealZoo | Embodied AI (Tracking, Navigation, Multi Agent) | Simulator | Website |
涵盖 RL 对齐、监督微调、相关开源仓库与提示工程。
| 标题 | 年份 | 论文 | RL 方法 | 代码 |
|---|---|---|---|---|
| Vero: An Open RL Recipe for General Visual Reasoning | 04/2026 | Paper | Task-routed rewards; GRPO-based | Code |
| wDPO: Winsorized Direct Preference Optimization for Robust Alignment | 03/2026 | Paper | wDPO | - |
| f-GRPO and Beyond: Divergence-Based RL for General LLM Alignment | 02/2026 | Paper | f-GRPO / f-HAL | |
| From Sight to Insight: Improving Visual Reasoning of MLLMs via Reinforcement Learning | 01/2026 | Paper | GRPO (6 reward functions) | |
| SaFeR-VLM: Safety-Aware Reinforcement Learning for Multimodal Reasoning | 2026 (ICLR) | Paper | GRPO + safety reward | |
| SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning | 11/2025 | Paper | Dual-Reward (Thinking + Judging) | |
| GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA | 10/2025 | Paper | GIFT (convex MSE loss) | |
| Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning | 10/12/2025 | Paper | GRPO | |
| Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play | 09/29/2025 | Paper | GRPO | - |
| Vision-SR1: Self-rewarding vision-language model via reasoning decomposition | 08/26/2025 | Paper | GRPO | - |
| Group Sequence Policy Optimization | 06/24/2025 | Paper | GSPO | - |
| Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning | 05/20/2025 | Paper | GRPO | - |
| VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | 2025/04/10 | Paper | GRPO | Code |
| OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement | 2025/03/21 | Paper | GRPO | Code |
| Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning | 2025/03/10 | Paper | GRPO | Code |
| OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference | 2025 | Paper | DPO | Code |
| Multimodal Open R1/R1-Multimodal-Journey | 2025 | - | GRPO | Code |
| R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization | 2025 | Paper | GRPO | Code |
| Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning | 2025 | - | PPO/REINFORCE++/GRPO | Code |
| MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement Learning | 2025 | Paper | REINFORCE Leave-One-Out (RLOO) | Code |
| MM-RLHF: The Next Step Forward in Multimodal LLM Alignment | 2025 | Paper | DPO | Code |
| LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL | 2025 | Paper | PPO | Code |
| Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models | 2025 | Paper | GRPO | Code |
| Unified Reward Model for Multimodal Understanding and Generation | 2025 | Paper | DPO | Code |
| Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step | 2025 | Paper | DPO | Code |
| All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning | 2025 | Paper | Online RL | - |
| Video-R1: Reinforcing Video Reasoning in MLLMs | 2025 | Paper | GRPO | Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes | 08/2026 | Paper | - | - |
| SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them | 07/2026 | Paper | - | - |
| AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of VLMs | 2026/03 | Paper | - | - |
| CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models | 2026/03 | Paper | - | |
| MERGETUNE: Continued Fine-Tuning of Vision-Language Models | 2026/01 (ICLR 2026) | Paper | - | |
| Mask Fine-Tuning (MFT): Unlocking Hidden Capabilities in Vision-Language Models | 2025/12 | Paper | - | |
| Image-LoRA: Towards Minimal Fine-Tuning of VLMs | 2025/12 | Paper | - | |
| Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning | 2025/12 | Paper | - | |
| Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models | 2025/04/21 | Paper | Website | |
| OMNICAPTIONER: One Captioner to Rule Them All | 2025/04/09 | Paper | Website | Code |
| Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning | 2024 | Paper | Website | Code |
| LLaVolta: Efficient Multi-modal Models via Stage-wise Visual Context Compression | 2024 | Paper | Website | Code |
| ViTamin: Designing Scalable Vision Models in the Vision-Language Era | 2024 | Paper | Website | Code |
| Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model | 2024 | Paper | - | - |
| Should VLMs be Pre-trained with Image Data? | 2025 | Paper | - | - |
| VisionArena: 230K Real World User-VLM Conversations with Preference Labels | 2024 | Paper | - | Code |
| 项目 | 仓库链接 |
|---|---|
| Verl | 🔗 GitHub |
| EasyR1 | 🔗 GitHub |
| OpenR1 | 🔗 GitHub |
| LLaMAFactory | 🔗 GitHub |
| MM-Eureka-Zero | 🔗 GitHub |
| MM-RLHF | 🔗 GitHub |
| LMM-R1 | 🔗 GitHub |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| EvoPrompt: Evolving Prompt Adaptation for Vision-Language Models | 2026/03 | Paper | - | - |
| MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation | 2026/02 | Paper | - | |
| Multimodal Prompt Optimizer (MPO): Joint Optimization of Multimodal Prompts | 2025/10 | Paper | - | |
| Evolutionary Prompt Optimization Discovers Emergent Multimodal Reasoning Strategies | 2025/03 | Paper | - | |
| In-ContextEdit:EnablingInstructionalImageEditingwithIn-Context GenerationinLargeScaleDiffusionTransformer | 2025/04/30 | Paper | Website |
| 标题 | 年份 | 论文链接 |
|---|---|---|
| Metis: Memory Foundation Model | 07/2026 | Paper |
| Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI | 2024 | Paper |
| ScreenAI: A Vision-Language Model for UI and Infographics Understanding | 2024 | Paper |
| ChartLlama: A Multimodal LLM for Chart Understanding and Generation | 2023 | Paper |
| SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement | 2024 | 📄 Paper |
| Training a Vision Language Model as Smartphone Assistant | 2024 | Paper |
| ScreenAgent: A Vision-Language Model-Driven Computer Control Agent | 2024 | Paper |
| Embodied Vision-Language Programmer from Environmental Feedback | 2024 | Paper |
| VLMs Play StarCraft II: A Benchmark and Multimodal Decision Method | 2025 | 📄 Paper |
| MP-GUI: Modality Perception with MLLMs for GUI Understanding | 2025 | 📄 Paper |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers | 07/2026 | 📄 Paper | - | - |
| GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning | 2023 | 📄 Paper | 🌍 Website | 💾 Code |
| Spurious Correlation in Multimodal LLMs | 2025 | 📄 Paper | - | - |
| WeGen: A Unified Model for Interactive Multimodal Generation as We Chat | 2025 | 📄 Paper | - | 💾 Code |
| VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation | 2024 | 📄 Paper | 🌍 Website | - |
| SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | 2024 | 📄 Paper | 🌍 Website | - |
| Vision-language model-driven scene understanding and robotic object manipulation | 2024 | 📄 Paper | - | - |
| Guiding Long-Horizon Task and Motion Planning with Vision Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| AutoTAMP: Autoregressive Task and Motion Planning with LLMs as Translators and Checkers | 2023 | 📄 Paper | 🌍 Website | - |
| VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model | 2024 | 📄 Paper | - | - |
| Scalable Multi-Robot Collaboration with Large Language Models: Centralized or Decentralized Systems? | 2023 | 📄 Paper | 🌍 Website | - |
| DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| MotionGPT: Human Motion as a Foreign Language | 2023 | 📄 Paper | - | 💾 Code |
| Learning Reward for Robot Skills Using Large Language Models via Self-Alignment | 2024 | 📄 Paper | - | - |
| Language to Rewards for Robotic Skill Synthesis | 2023 | 📄 Paper | 🌍 Website | - |
| Eureka: Human-Level Reward Design via Coding Large Language Models | 2023 | 📄 Paper | 🌍 Website | - |
| Integrated Task and Motion Planning | 2020 | 📄 Paper | - | - |
| Jailbreaking LLM-Controlled Robots | 2024 | 📄 Paper | 🌍 Website | - |
| Robots Enact Malignant Stereotypes | 2022 | 📄 Paper | 🌍 Website | - |
| LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions | 2024 | 📄 Paper | - | - |
| Highlighting the Safety Concerns of Deploying LLMs/VLMs in Robotics | 2024 | 📄 Paper | 🌍 Website | - |
| EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents | 2025 | 📄 Paper | 🌍 Website | 💾 Code & Dataset |
| Gemini Robotics: Bringing AI into the Physical World | 2025 | 📄 Technical Report | 🌍 Website | - |
| GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation | 2024 | 📄 Paper | 🌍 Website | - |
| Magma: A Foundation Model for Multimodal AI Agents | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| DayDreamer: World Models for Physical Robot Learning | 2022 | 📄 Paper | 🌍 Website | 💾 Code |
| Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models | 2025 | 📄 Paper | - | - |
| RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| Unified Video Action Model | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation | 03/2026 | 📄 Paper | - | |
| NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models | 03/2026 | 📄 Paper | - | |
| Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control | 02/2026 | 📄 Paper | - | |
| ST4VLA: Spatial Guided Training for Vision-Language-Action Models | 02/2026 | 📄 Paper | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data | 08/03/2026 | 📄 Paper | - | - |
| N₀-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens | 07/26/2026 | 📄 Paper | - | - |
| VIMA: General Robot Manipulation with Multimodal Prompts | 2022 | 📄 Paper | 🌍 Website | |
| Instruct2Act: Mapping Multi-Modality Instructions to Robotic Actions with Large Language Model | 2023 | 📄 Paper | - | - |
| Creative Robot Tool Use with Large Language Models | 2023 | 📄 Paper | 🌍 Website | - |
| RoboVQA: Multimodal Long-Horizon Reasoning for Robotics | 2024 | 📄 Paper | - | - |
| RT-1: Robotics Transformer for Real-World Control at Scale | 2022 | 📄 Paper | 🌍 Website | - |
| RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | 2023 | 📄 Paper | 🌍 Website | - |
| Open X-Embodiment: Robotic Learning Datasets and RT-X Models | 2023 | 📄 Paper | 🌍 Website | - |
| ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| Masked World Models for Visual Control | 2022 | 📄 Paper | 🌍 Website | 💾 Code |
| Multi-View Masked World Models for Visual Robotic Manipulation | 2023 | 📄 Paper | 🌍 Website | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings | 2022 | 📄 Paper | - | - |
| LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation | 2024 | 📄 Paper | - | - |
| LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action | 2022 | 📄 Paper | 🌍 Website | - |
| NaVILA: Legged Robot Vision-Language-Action Model for Navigation | 2022 | 📄 Paper | 🌍 Website | - |
| VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation | 2024 | 📄 Paper | - | - |
| Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning | 2023 | 📄 Paper | 🌍 Website | - |
| Vi-LAD: Vision-Language Attention Distillation for Socially-Aware Robot Navigation in Dynamic Environments | 2025 | 📄 Paper | - | - |
| Navigation World Models | 2024 | 📄 Paper | 🌍 Website | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| MUTEX: Learning Unified Policies from Multimodal Task Specifications | 2023 | 📄 Paper | 🌍 Website | - |
| LaMI: Large Language Models for Multi-Modal Human-Robot Interaction | 2024 | 📄 Paper | 🌍 Website | - |
| VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models | 2024 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving | 04/2026 | 📄 Paper | - | - |
| AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving | 03/2026 | 📄 Paper | - | |
| DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe Autonomous Driving | 03/2026 | 📄 Paper | - | |
| HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving | 02/2026 | 📄 Paper | - | |
| OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model | 03/2025 | 📄 Paper | - | |
| Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives | 01/07/2025 | 📄 Paper | 🌍 Website | |
| DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| GPT-Driver: Learning to Drive with GPT | 2023 | 📄 Paper | - | - |
| LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving | 2023 | 📄 Paper | 🌍 Website | - |
| Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving | 2023 | 📄 Paper | - | - |
| Referring Multi-Object Tracking | 2023 | 📄 Paper | - | 💾 Code |
| VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision | 2023 | 📄 Paper | - | 💾 Code |
| MotionLM: Multi-Agent Motion Forecasting as Language Modeling | 2023 | 📄 Paper | - | - |
| DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models | 2023 | 📄 Paper | 🌍 Website | - |
| VLP: Vision Language Planning for Autonomous Driving | 2024 | 📄 Paper | - | - |
| DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model | 2023 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis | 2024 | 📄 Paper | - | 💾 Code |
| LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration – A Robot Sous-Chef Application | 2024 | 📄 Paper | - | - |
| Pretrained Language Models as Visual Planners for Human Assistance | 2023 | 📄 Paper | - | - |
| Promoting AI Equity in Science: Generalized Domain Prompt Learning for Accessible VLM Research | 2024 | 📄 Paper | - | - |
| Image and Data Mining in Reticular Chemistry Using GPT-4V | 2023 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent | 08/05/2026 | 📄 Paper | - | - |
| StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents | 07/24/2026 | 📄 Paper | - | - |
| A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis | 2023 | 📄 Paper | - | - |
| CogAgent: A Visual Language Model for GUI Agents | 2023 | 📄 Paper | - | 💾 Code |
| WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models | 2024 | 📄 Paper | - | 💾 Code |
| ShowUI: One Vision-Language-Action Model for GUI Visual Agent | 2024 | 📄 Paper | - | 💾 Code |
| ScreenAgent: A Vision Language Model-driven Computer Control Agent | 2024 | 📄 Paper | - | 💾 Code |
| Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation | 2024 | 📄 Paper | - | 💾 Code |
| MolmoWeb: Open Visual Web Agent and Open Data for the Open Web | 04/2026 | 📄 Paper | 🌍 Website |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| X-World: Accessibility, Vision, and Autonomy Meet | 2021 | 📄 Paper | - | - |
| Context-Aware Image Descriptions for Web Accessibility | 2024 | 📄 Paper | - | - |
| Improving VR Accessibility Through Automatic 360 Scene Description Using Multimodal Large Language Models | 2024 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| RESPClinBench: Multimodal Clinical Decision-Making and Longitudinal Disease Tracking | 08/07/2026 | 📄 Paper | - | - |
| CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework | 03/2026 | 📄 Paper | - | - |
| MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images | 02/2026 | 📄 Paper | - | |
| Colon-X: Advancing Intelligent Colonoscopy from Multimodal Understanding to Clinical Reasoning | 12/2025 | 📄 Paper | - | |
| Frontiers in Intelligent Colonoscopy | 02/2025 | 📄 Paper | - | 💾 Code |
| VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge | 2024 | 📄 Paper | - | 💾 Code |
| Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for Radiology | 2024 | 📄 Paper | - | - |
| M-FLAG: Medical Vision-Language Pre-training with Frozen Language Models and Latent Space Geometry Optimization | 2023 | 📄 Paper | - | - |
| MedCLIP: Contrastive Learning from Unpaired Medical Images and Text | 2022 | 📄 Paper | - | 💾 Code |
| Med-Flamingo: A Multimodal Medical Few-Shot Learner | 2023 | 📄 Paper | - | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Analyzing K-12 AI Education: A Large Language Model Study of Classroom Instruction on Learning Theories, Pedagogy, Tools, and AI Literacy | 2024 | 📄 Paper | - | - |
| Students Rather Than Experts: A New AI for Education Pipeline to Model More Human-Like and Personalized Early Adolescence | 2024 | 📄 Paper | - | - |
| Harnessing Large Vision and Language Models in Agriculture: A Review | 2024 | 📄 Paper | - | - |
| A Vision-Language Model for Predicting Potential Distribution Land of Soybean Double Cropping | 2024 | 📄 Paper | - | - |
| Vision-Language Model is NOT All You Need: Augmentation Strategies for Molecule Language Models | 2024 | 📄 Paper | - | 💾 Code |
| DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images | 2024 | 📄 Paper | - | - |
| MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models | 2024 | 📄 Paper | - | 💾 Code |
| Vision-Language Models Meet Meteorology: Developing Models for Extreme Weather Events Detection with Heatmaps | 2024 | 📄 Paper | - | 💾 Code |
| He is Very Intelligent, She is Very Beautiful? On Mitigating Social Biases in Language Modeling and Generation | 2021 | 📄 Paper | - | - |
| UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Region Profiling | 2024 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs | 08/07/2026 | 📄 Paper | - | - |
| Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models | 04/2026 | 📄 Paper | - | - |
| VLMs Need Words: Vision Language Models Ignore Visual Detail in Favor of Semantic Anchors | 04/2026 | 📄 Paper | - | |
| HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token | 03/2026 | 📄 Paper | 🌍 ACL | |
| Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs | 01/2026 | 📄 Paper | - | |
| Object Hallucination in Image Captioning | 2018 | 📄 Paper | - | |
| Evaluating Object Hallucination in Large Vision-Language Models | 2023 | 📄 Paper | - | 💾 Code |
| Detecting and Preventing Hallucinations in Large Vision Language Models | 2023 | 📄 Paper | - | - |
| HallE-Control: Controlling Object Hallucination in Large Multimodal Models | 2023 | 📄 Paper | - | 💾 Code |
| Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs | 2024 | 📄 Paper | - | 💾 Code |
| BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models | 2023 | 📄 Paper | - | 💾 Code |
| AUTOHALLUSION: Automatic Generation of Hallucination Benchmarks for Vision-Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning | 2023 | 📄 Paper | - | 💾 Code |
| Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models | 2024 | 📄 Paper | - | 💾 Code |
| AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation | 2023 | 📄 Paper | - | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| SaFeR-VLM: Safety into Multimodal Reasoning via Reinforcement Learning | 2026 (ICLR) | 📄 Paper | - | - |
| HoliSafe: Holistic Safety Evaluation for Vision-Language Models | 2026 (ICLR) | 📄 Paper | - | |
| JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models | 2024 | 📄 Paper | 🌍 Website | |
| Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments | 2023 | 📄 Paper | - | - |
| SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models | 2024 | 📄 Paper | - | - |
| JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks | 2024 | 📄 Paper | - | - |
| SHIELD: An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models | 2024 | 📄 Paper | - | 💾 Code |
| Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models | 2024 | 📄 Paper | - | - |
| Jailbreaking Attack against Multimodal Large Language Model | 2024 | 📄 Paper | - | - |
| Embodied Red Teaming for Auditing Robotic Foundation Models | 2025 | 📄 Paper | 🌍 Website | |
| Safety Guardrails for LLM-Enabled Robots | 2025 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Hallucination of Multimodal Large Language Models: A Survey | 2024 | 📄 Paper | - | - |
| Bias and Fairness in Large Language Models: A Survey | 2023 | 📄 Paper | - | - |
| Fairness and Bias in Multimodal AI: A Survey | 2024 | 📄 Paper | - | - |
| Multi-Modal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision–Language Models | 2023 | 📄 Paper | - | - |
| FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks | 2024 | 📄 Paper | - | - |
| FairCLIP: Harnessing Fairness in Vision-Language Learning | 2024 | 📄 Paper | - | - |
| FairMedFM: Fairness Benchmarking for Medical Imaging Foundation Models | 2024 | 📄 Paper | - | - |
| Benchmarking Vision Language Models for Cultural Understanding | 2024 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding | 2024 | 📄 Paper | - | - |
| Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement | 2024 | 📄 Paper | - | - |
| Assessing and Learning Alignment of Unimodal Vision and Language Models | 2024 | 📄 Paper | 🌍 Website | - |
| Extending Multi-modal Contrastive Representations | 2023 | 📄 Paper | - | 💾 Code |
| OneLLM: One Framework to Align All Modalities with Language | 2023 | 📄 Paper | - | 💾 Code |
| What You See is What You Read? Improving Text-Image Alignment Evaluation | 2023 | 📄 Paper | 🌍 Website | 💾 Code |
| Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| VBench: Comprehensive BenchmarkSuite for Video Generative Models | 2023 | 📄 Paper | 🌍 Website | 💾 Code |
| VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| PhysBench: Benchmarking and Enhancing VLMs for Physical World Understanding | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| VideoPhy: Evaluating Physical Commonsense for Video Generation | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| WorldSimBench: Towards Video Generation Models as World Simulators | 2024 | 📄 Paper | 🌍 Website | - |
| WorldModelBench: Judging Video Generation Models As World Models | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation | 2025 | 📄 Paper | - | 💾 Code |
| Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency | 2025 | 📄 Paper | - | 💾 Code |
| Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding | 2025 | 📄 Paper | - | - |
| SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| Do generative video models understand physical principles? | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| How Far is Video Generation from World Model: A Physical Law Perspective | 2024 | 📄 Paper | 🌍 Website | 💾 Code |
| Imagine while Reasoning in Space: Multimodal Visualization-of-Thought | 2025 | 📄 Paper | - | - |
| VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness | 2025 | 📄 Paper | 🌍 Website | 💾 Code |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models | 08/05/2026 | 📄 Paper | - | - |
| GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video LLMs | 08/05/2026 | 📄 Paper | - | - |
| QAPruner: Quantization-Aware Vision Token Pruning for MLLMs | 04/2026 | 📄 Paper | - | - |
| Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation | 04/2026 | 📄 Paper | - | |
| CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning | 04/2026 | 📄 Paper | - | |
| LoRA-Squeeze: Simple and Effective Post-Tuning and In-Tuning Compression of LoRA Modules | 02/2026 | 📄 Paper | - | |
| GRACE: Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs | 01/2026 | 📄 Paper | - | |
| VLMQ: Post-Training Quantization for Large Vision-Language Models | 2026 (ICLR) | 📄 Paper | - | |
| VILA: On Pre-training for Visual Language Models | 2023 | 📄 Paper | - | |
| SimVLM: Simple Visual Language Model Pretraining with Weak Supervision | 2021 | 📄 Paper | - | - |
| LoRA: Low-Rank Adaptation of Large Language Models | 2021 | 📄 Paper | - | 💾 Code |
| QLoRA: Efficient Finetuning of Quantized LLMs | 2023 | 📄 Paper | - | - |
| Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback | 2022 | 📄 Paper | - | 💾 Code |
| RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback | 2023 | 📄 Paper | - | - |
| 标题 | 年份 | 论文 | 站点 | 代码 |
|---|---|---|---|---|
| A Survey on Bridging VLMs and Synthetic Data | 2025 | 📄 Paper | - | 💾 Code |
| Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning | 2024 | 📄 Paper | Website | 💾 Code |
| SLIP: Self-supervision meets Language-Image Pre-training | 2021 | 📄 Paper | - | 💾 Code |
| Synthetic Vision: Training Vision-Language Models to Understand Physics | 2024 | 📄 Paper | - | - |
| Synth2: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings | 2024 | 📄 Paper | - | - |
| KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data | 2024 | 📄 Paper | - | - |
| Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation | 2024 | 📄 Paper | - | - |