DEVIS：使变形变压器用于视频实例细分

论文标题

DEVIS：使变形变压器用于视频实例细分

DeVIS: Making Deformable Transformers Work for Video Instance Segmentation

论文作者

Caelles, Adrià, Meinhardt, Tim, Brasó, Guillem, Leal-Taixé, Laura

论文摘要

视频实例分割（VIS）在视频序列中共同处理多对象检测，跟踪和分割。过去，VIS方法反映了这些子任务在其建筑设计中的碎片化，因此在关节溶液上错过了这些子任务。变形金刚最近允许将整个VIS任务作为单个设定预测问题进行。然而，现有基于变压器的方法的二次复杂性需要较长的训练时间，高内存需求以及低音尺度特征图的处理。可变形的注意力提供了更有效的替代方案，但尚未探索其对时间域或分割任务的应用。在这项工作中，我们提出了可变形的Vis（Devis），这是一种利用可变形变压器的效率和性能的VIS方法。为了在多个帧上共同考虑所有VIS子任务，我们使用实例感知对象查询表示时间尺度可变形。我们进一步介绍了带有多尺度功能的新图像和视频实例蒙版头，并使用多提示剪辑跟踪执行近线视频处理。 Devis减少了内存和训练时间要求，并在YouTube-Vis 2021以及具有挑战性的OVIS数据集中实现了最新的结果。代码可从https://github.com/acaelles97/devis获得。

Video Instance Segmentation (VIS) jointly tackles multi-object detection, tracking, and segmentation in video sequences. In the past, VIS methods mirrored the fragmentation of these subtasks in their architectural design, hence missing out on a joint solution. Transformers recently allowed to cast the entire VIS task as a single set-prediction problem. Nevertheless, the quadratic complexity of existing Transformer-based methods requires long training times, high memory requirements, and processing of low-single-scale feature maps. Deformable attention provides a more efficient alternative but its application to the temporal domain or the segmentation task have not yet been explored. In this work, we present Deformable VIS (DeVIS), a VIS method which capitalizes on the efficiency and performance of deformable Transformers. To reason about all VIS subtasks jointly over multiple frames, we present temporal multi-scale deformable attention with instance-aware object queries. We further introduce a new image and video instance mask head with multi-scale features, and perform near-online video processing with multi-cue clip tracking. DeVIS reduces memory as well as training time requirements, and achieves state-of-the-art results on the YouTube-VIS 2021, as well as the challenging OVIS dataset. Code is available at https://github.com/acaelles97/DeVIS.

下载PDF全文

下载文献需遵守相关版权规定

论文标题