학술논문

Hybrid CNN and Transformer Network for Semantic Segmentation of UAV Remote Sensing Images
Document Type
Periodical
Source
IEEE Journal on Miniaturization for Air and Space Systems IEEE J. Miniat. Air Space Syst. Miniaturization for Air and Space Systems, IEEE Journal on. 5(1):33-41 Mar, 2024
Subject
Aerospace
Components, Circuits, Devices and Systems
Transportation
Remote sensing
Autonomous aerial vehicles
Transformers
Semantic segmentation
Feature extraction
Convolutional neural networks
semantic segmentation
Swin transformer
unmanned aerial vehicle (UAV)
Language
ISSN
2576-3164
Abstract
Semantic segmentation of unmanned aerial vehicle (UAV) remote sensing images is a recent research hotspot, offering technical support for diverse types of UAV remote sensing missions. However, unlike general scene images, UAV remote sensing images present inherent challenges. These challenges include the complexity of backgrounds, substantial variations in target scales, and dense arrangements of small targets, which severely hinder the accuracy of semantic segmentation. To address these issues, we propose a convolutional neural network (CNN) and transformer hybrid network for semantic segmentation of UAV remote sensing images. The proposed network follows an encoder–decoder architecture that merges a transformer-based encoder with a CNN-based decoder. First, we incorporate the Swin transformer as the encoder to address the limitations of CNN in global modeling, mitigating the interference caused by complex background information. Second, to effectively handle the significant changes in target scales, we design the multiscale feature integration module (MFIM) that enhances the multiscale feature representation capability of the network. Finally, the semantic feature fusion module (SFFM) is designed to filter the redundant noise during the feature fusion process, which improves the recognition of small targets and edges. Experimental results demonstrate that the proposed method outperforms other popular methods on the UAVid and Aeroscapes datasets.