Positive Sample Propagation along the Audio-Visual Event Line
Zhou, Zheng, Zhong, Hao, Wang · cs.CV,cs.MM,cs.SD,eess.AS · 2026-09-05 · 原文
Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or positive) audio-visual segment pairs while filtering out the irrelevant ones, regardless whether they are synchronized or not. To this end, we propose a new positive sample propagation (PSP) module to discover and exploit the closely related audio-visual pairs by evaluating the relationship within every possible pair. It can be done by constructing an all-pair similarity map between each audio and visual segment, and only aggregating the features from the pairs with high similarity scores. To encourage the network to extract high correlated features for positive samples, a new audio-visual pair similarity loss is proposed. We also propose a new weighting branch to better exploit the temporal correlations in weakly supervised setting. We perform extensive experiments on the public AVE dataset and achieve new state-of-the-art accuracy in both fully and weakly supervised settings, thus verifyin
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs).
Given a video, we aim to localize video segments containing an AVE and identify its category.
2. 领域脉络
本文类目:cs.CV、cs.MM、cs.SD、eess.AS,属于其所在研究脉络的最新进展。
3. 机制拆解
In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or positive) audio-visual segment pairs while filtering out the irrelevant ones, regardless whether they are synchronized or not.
To this end, we propose a new positive sample propagation (PSP) module to discover and exploit the closely related audio-visual pairs by evaluating the relationship within every possible pair.
4. 证据与数字
摘要未给出量化结果——留意原文的实验与数据。
5. 反例与边界
It can be done by constructing an all-pair similarity map between each audio and visual segment, and only aggregating the features from the pairs with high similarity scores.
6. 跨领域连接与意外收获
横跨 4 个类目(cs.CV、cs.MM、cs.SD、eess.AS),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。