高级检索

    面向遥感图像建筑物提取的轻型注意力网络

    A lightweight attention network for feature extraction of buildings from remote sensing images

    • 摘要: 为解决高分辨率遥感图像建筑物提取任务中普遍存在的分割精度低与计算资源消耗大的问题,该文提出了一种基于双边感知网络(bilateral awareness network, BANet)的轻型注意力网络(lightweight attention network,LANet)。首先,为了充分捕获丰富的全局语义信息,采用BuildFormer骨干网络作为全局特征提取分支; 其次,为了实现跨尺度的上下文建模与长程依赖捕捉,引入多尺度线性注意力(multi-scale linear attention,MSLA)机制以有效捕获重要信息; 最后,为了提升模型对感兴趣区域的关注度并抑制无关背景干扰,引入通道优先卷积注意力(channel prior convolutional attention,CPCA)机制。为了验证该文提出的LANet的有效性与适用性,在开源的WHU和Massachusetts建筑数据集上开展了实验, 结果表明: LANet在精度与效率方面均表现出显著优势, 具体而言,该模型参数量和计算量仅为17.11×106个和18.69×109次, 在2个测试集上的交并比分别达到了91.22%和75.72%。相较于一些先进的模型,LANet表现出了优越的综合性能,表明其能够在计算资源受限的场景中准确地提取建筑物特征。

       

      Abstract: To address the low segmentation accuracy and high computational resource consumption commonly encountered in feature extraction of buildings from high-resolution remote sensing images, this study developed a BANet-based lightweight attention network (LANet). First, to fully capture rich global semantic information, the BuildFormer backbone network was adopted as the branch for global feature extraction. Second, to enable cross-scale contextual modeling and long-range dependency capture, a multi-scale linear attention (MSLA) mechanism was introduced to effectively capture important information. Third, a channel prior convolutional attention (CPCA) mechanism was incorporated to enhance the model's focus on regions of interest while suppressing irrelevant background interference. Finally, the effectiveness and applicability of the proposed LANet were validated on the open-source WHU and Massachusetts building datasets. The experimental results demonstrate that the LANet offered significant advantages in terms of both segmentation accuracy and computational efficiency. Specifically, the LANet yielded intersection over union (IoU) values of 91.22% and 75.72% on the two test sets, requiring merely 17.11×106 of parameters and 18.69×109 of computational load. Compared to some advanced models, the LANet demonstrates superior comprehensive performance, suggesting that it has the capability to accurately extract the features of buildings under limited computational resources.

       

    /

    返回文章
    返回