Abstract:
To address the low segmentation accuracy and high computational resource consumption commonly encountered in feature extraction of buildings from high-resolution remote sensing images, this study developed a BANet-based lightweight attention network (LANet). First, to fully capture rich global semantic information, the BuildFormer backbone network was adopted as the branch for global feature extraction. Second, to enable cross-scale contextual modeling and long-range dependency capture, a multi-scale linear attention (MSLA) mechanism was introduced to effectively capture important information. Third, a channel prior convolutional attention (CPCA) mechanism was incorporated to enhance the model's focus on regions of interest while suppressing irrelevant background interference. Finally, the effectiveness and applicability of the proposed LANet were validated on the open-source WHU and Massachusetts building datasets. The experimental results demonstrate that the LANet offered significant advantages in terms of both segmentation accuracy and computational efficiency. Specifically, the LANet yielded intersection over union (IoU) values of 91.22% and 75.72% on the two test sets, requiring merely 17.11×10
6 of parameters and 18.69×10
9 of computational load. Compared to some advanced models, the LANet demonstrates superior comprehensive performance, suggesting that it has the capability to accurately extract the features of buildings under limited computational resources.