解决视频对象细分中的背景干扰

论文标题

解决视频对象细分中的背景干扰

Tackling Background Distraction in Video Object Segmentation

论文作者

Cho, Suhwan, Lee, Heansung, Lee, Minhyeok, Park, Chaewon, Jang, Sungjun, Kim, Minjung, Lee, Sangyoun

论文摘要

半监督视频对象细分（VOS）旨在密集跟踪视频中的某些指定对象。该任务中的主要挑战之一是存在与目标对象相似的背景干扰器的存在。我们提出了三种抑制此类干扰因素的新型策略：1）一种时空多元化的模板构建方案，以获得目标对象的广义特性； 2）可学习的距离得分函数，可通过利用两个连续帧之间的时间一致性来排除空间距离的干扰因素； 3）交换和连接的增强通过提供包含纠缠对象的训练样本，迫使每个对象具有独特的功能。在所有公共基准数据集中，即使是实时性能，我们的模型也与当代最先进的方法相当。定性结果还证明了我们的方法优于现有方法。我们认为，我们的方法将被广泛用于未来的VOS研究。

Semi-supervised video object segmentation (VOS) aims to densely track certain designated objects in videos. One of the main challenges in this task is the existence of background distractors that appear similar to the target objects. We propose three novel strategies to suppress such distractors: 1) a spatio-temporally diversified template construction scheme to obtain generalized properties of the target objects; 2) a learnable distance-scoring function to exclude spatially-distant distractors by exploiting the temporal consistency between two consecutive frames; 3) swap-and-attach augmentation to force each object to have unique features by providing training samples containing entangled objects. On all public benchmark datasets, our model achieves a comparable performance to contemporary state-of-the-art approaches, even with real-time performance. Qualitative results also demonstrate the superiority of our approach over existing methods. We believe our approach will be widely used for future VOS research.

下载PDF全文

下载文献需遵守相关版权规定

论文标题