Abstract

Image segmentation, the task of assigning a class label to each pixel in an image, has been widely applied in remote sensing to identify surficial geological and hydrological features from data sources such as multispectral satellite imagery, aerial photography, and digital elevation models. UNet is among the most widely adopted convolutional neural network architectures for segmentation, owing to its encoder–decoder structure and skip connections that enable precise spatial localization. UNet has been applied to detect sinkholes using lidar-derived high-resolution digital elevation data; however, existing models still fall short of the accuracy of manual sinkhole mapping. In this study, we aim to improve UNet-based sinkhole segmentation by incorporating attention mechanisms and fusing elevation data with aerial imagery. Attention mechanisms allow the model to learn to focus on the most task-relevant parts of the input, selectively emphasizing features and spatial regions indicative of sinkholes. Aerial imagery provides complementary visual cues—such as vegetation changes near or within sinkholes—that are routinely leveraged in manual mapping, and fusing it with elevation data may help distinguish sinkholes from similar topographic features. We trained and evaluated our models using high-resolution lidar-derived elevation data, plane-based leaf-off aerial imagery, and high-accuracy manually mapped sinkhole inventories from multiple areas in Kentucky. We compared the performance of UNet and an attention-enhanced UNet, referred to as UNet-T, across different combinations of input modalities. Our results suggest that model generalization improves with increasing number of training areas, and the combination of DEM gradients and topographic derivatives (slope, aspect, and curvature) yields the strongest terrain representation. In addition, fusion of aerial imagery with DEM inputs can enhance model performance. Conversely, UNet-T provides only a marginal and inconsistent advantage over UNet.

DOI

https://doi.org/10.5038/9781967518012.1006

Zhu_Table1.pdf (80 kB)
Table 1

Zhu_Table2.pdf (49 kB)
Table 2

Zhu_Fig.1.tif (323 kB)
Figure 1

Zhu_Fig.2.tif (1459 kB)
Figure 2

Zhu_Fig.3.tif (13650 kB)
Figure 3

Zhu_Fig.4.tif (17937 kB)
Figure 4

Zhu_Fig.5.tif (21016 kB)
Figure 5

Zhu_Fig.6.tif (26430 kB)
Figure 6

Responses_to_reviewer_comments.docx (26 kB)
Response to reviewer comments

Share

COinS
 

Attention-Enhanced Multimodal Fusion of Elevation and Aerial Imagery for Sinkhole Segmentation

Image segmentation, the task of assigning a class label to each pixel in an image, has been widely applied in remote sensing to identify surficial geological and hydrological features from data sources such as multispectral satellite imagery, aerial photography, and digital elevation models. UNet is among the most widely adopted convolutional neural network architectures for segmentation, owing to its encoder–decoder structure and skip connections that enable precise spatial localization. UNet has been applied to detect sinkholes using lidar-derived high-resolution digital elevation data; however, existing models still fall short of the accuracy of manual sinkhole mapping. In this study, we aim to improve UNet-based sinkhole segmentation by incorporating attention mechanisms and fusing elevation data with aerial imagery. Attention mechanisms allow the model to learn to focus on the most task-relevant parts of the input, selectively emphasizing features and spatial regions indicative of sinkholes. Aerial imagery provides complementary visual cues—such as vegetation changes near or within sinkholes—that are routinely leveraged in manual mapping, and fusing it with elevation data may help distinguish sinkholes from similar topographic features. We trained and evaluated our models using high-resolution lidar-derived elevation data, plane-based leaf-off aerial imagery, and high-accuracy manually mapped sinkhole inventories from multiple areas in Kentucky. We compared the performance of UNet and an attention-enhanced UNet, referred to as UNet-T, across different combinations of input modalities. Our results suggest that model generalization improves with increasing number of training areas, and the combination of DEM gradients and topographic derivatives (slope, aspect, and curvature) yields the strongest terrain representation. In addition, fusion of aerial imagery with DEM inputs can enhance model performance. Conversely, UNet-T provides only a marginal and inconsistent advantage over UNet.