EdgeTAM: On-Device Track anything Model

페이지 정보

profile_image
작성자 Hope
댓글 0건 조회 38회 작성일 25-09-21 04:47

본문

On high of Segment Anything Model (SAM), SAM 2 additional extends its capability from picture to video inputs by a reminiscence bank mechanism and obtains a remarkable efficiency compared with previous methods, making it a basis mannequin for video segmentation job. In this paper, we aim at making SAM 2 way more environment friendly so that it even runs on cellular units whereas maintaining a comparable performance. Despite a number of works optimizing SAM for better efficiency, we find they are not ample for SAM 2 because they all concentrate on compressing the picture encoder, whereas our benchmark exhibits that the newly launched reminiscence attention blocks are also the latency bottleneck. Given this commentary, we propose EdgeTAM, which leverages a novel 2D Spatial Perceiver to scale back the computational value. In particular, the proposed 2D Spatial Perceiver encodes the densely stored body-stage reminiscences with a lightweight Transformer that comprises a set set of learnable queries.



65d43cc49290ca0a32180020_v2-7lhou-5cpai.jpegOn condition that video segmentation is a dense prediction task, we discover preserving the spatial construction of the memories is important so that the queries are cut up into international-level and patch-level groups. We also propose a distillation pipeline that further improves the performance without inference overhead. DAVIS 2017, MOSE, SA-V val, and SA-V test, while running at 16 FPS on iPhone 15 Pro Max. SAM to handle each image and video inputs, with a memory financial institution mechanism, and is skilled with a brand iTagPro USA new large-scale multi-grained video tracking dataset (SA-V). Despite attaining an astonishing efficiency in comparison with earlier video object segmentation (VOS) models and permitting extra numerous person prompts, iTagPro USA SAM 2, as a server-facet basis model, is just not efficient for on-device inference. CPU and NPU. Throughout the paper, we interchangeably use iPhone and iPhone 15 Pro Max for simplicity.. SAM for better efficiency only consider squeezing its picture encoder since the mask decoder is extraordinarily lightweight. SAM 2. Specifically, SAM 2 encodes previous frames with a memory encoder, and these frame-degree recollections together with object-level pointers (obtained from the mask decoder) serve as the memory bank.



These are then fused with the options of present frame via memory attention blocks. As these reminiscences are densely encoded, this results in a huge matrix multiplication in the course of the cross-attention between present body features and ItagPro memory options. Therefore, regardless of containing relatively fewer parameters than the image encoder, the computational complexity of the reminiscence consideration just isn't reasonably priced for on-machine inference. The speculation is further proved by Fig. 2, iTagPro USA where lowering the number of reminiscence attention blocks almost linearly cuts down the general decoding latency and within each memory consideration block, removing the cross attention provides the most important pace-up. To make such a video-primarily based monitoring model run on gadget, iTagPro technology in EdgeTAM, we look at exploiting the redundancy in videos. To do this in observe, we propose to compress the uncooked body-degree memories earlier than performing memory consideration. We begin with naïve spatial pooling and observe a major performance degradation, particularly when using low-capability backbones.



However, naïvely incorporating a Perceiver also results in a extreme drop in efficiency. We hypothesize that as a dense prediction process, the video segmentation requires preserving the spatial structure of the reminiscence bank, which a naïve Perceiver discards. Given these observations, we propose a novel lightweight module that compresses body-degree memory feature maps whereas preserving the 2D spatial structure, iTagPro device named 2D Spatial Perceiver. Specifically, we split the learnable queries into two teams, the place one group features equally to the original Perceiver, the place every question performs international attention on the enter features and outputs a single vector as the body-stage summarization. In the opposite group, the queries have 2D priors, i.e., each query is simply chargeable for compressing a non-overlapping native patch, thus the output maintains the spatial structure while reducing the full variety of tokens. In addition to the structure improvement, we further propose a distillation pipeline that transfers the information of the powerful trainer SAM 2 to our scholar mannequin, which improves the accuracy for free of charge of inference overhead.



We find that in both stages, aligning the features from image encoders of the unique SAM 2 and our environment friendly variant advantages the performance. Besides, we further align the function output from the memory consideration between the trainer SAM 2 and our pupil mannequin within the second stage in order that in addition to the image encoder, memory-associated modules may also obtain supervision indicators from the SAM 2 teacher. SA-V val and ItagPro test by 1.Three and 3.3, respectively. Putting together, we propose EdgeTAM (Track Anything Model for Edge gadgets), iTagPro USA that adopts a 2D Spatial Perceiver for effectivity and knowledge distillation for accuracy. Through comprehensive benchmark, we reveal that the latency bottleneck lies in the reminiscence consideration module. Given the latency analysis, iTagPro USA we suggest a 2D Spatial Perceiver that considerably cuts down the reminiscence attention computational value with comparable performance, iTagPro USA which can be integrated with any SAM 2 variants. We experiment with a distillation pipeline that performs characteristic-smart alignment with the unique SAM 2 in each the picture and video segmentation levels and observe efficiency enhancements with none extra value throughout inference.

댓글목록

등록된 댓글이 없습니다.