ADP 前沿学习

← 板块一 · 研究前沿

ChatBEV: Empowering Traffic Scene Understanding and Simulation via Vision-Language Model

Xu, Zhang, Wang, Chen · cs.CV,cs.AI · 2026-09-05 · 原文

Comprehensive traffic scene understanding is a foundational capability for Intelligent Transportation Systems (ITS) underpinning applications such as traffic simulation. While VisionLanguage Models (VLMs) have demonstrated strong reasoning potential, their application to Bird's-Eye View (BEV) maps in traffic contexts remains limited by narrow task definitions and scarce annotated data. We introduce ChatBEV-QA, a large-scale BEV VQA benchmark of 137K+ QA pairs, designed to evaluate global scene understanding, vehicle-lane interactions, and vehiclevehicle interactions within complex traffic environments. Building on this, we fine-tune ChatBEV, a specialized VLM that accurately interprets diverse scene understanding queries from BEV maps. To demonstrate downstream utility in ITS applications, we integrate ChatBEV into a language-guided traffic simulation framework. Its global understanding and navigation reasoning provide crucial context-aware guidance, reducing trajectory displacement error by up to 20.9% and scenario collision rates by up to 37.9% over text-only baselines.

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

Comprehensive traffic scene understanding is a foundational capability for Intelligent Transportation Systems (ITS) underpinning applications such as traffic simulation.

While VisionLanguage Models (VLMs) have demonstrated strong reasoning potential, their application to Bird's-Eye View (BEV) maps in traffic contexts remains limited by narrow task definitions and scarce annotated data.

2. 领域脉络

本文类目:cs.CV、cs.AI,属于其所在研究脉络的最新进展。

3. 机制拆解

Building on this, we fine-tune ChatBEV, a specialized VLM that accurately interprets diverse scene understanding queries from BEV maps.

To demonstrate downstream utility in ITS applications, we integrate ChatBEV into a language-guided traffic simulation framework.

4. 证据与数字

We introduce ChatBEV-QA, a large-scale BEV VQA benchmark of 137K+ QA pairs, designed to evaluate global scene understanding, vehicle-lane interactions, and vehiclevehicle interactions within complex traffic environments.

Its global understanding and navigation reasoning provide crucial context-aware guidance, reducing trajectory displacement error by up to 20.9% and scenario collision rates by up to 37.9% over text-only baselines.

5. 反例与边界

摘要未声明局限与反例——这是需要警惕的信号,精读时先问边界。

6. 跨领域连接与意外收获

横跨 2 个类目(cs.CV、cs.AI),关注其在你兴趣板块间的迁移面。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。