S3CFSL: Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification

Mengxin Cao, Xu Zhang, Jinyong Cheng, Guixin Zhao, Wei Li, Xiangjun Dong

IEEE Transactions on Geoscience and Remote Sensing, vol. 62, article 5525315 (2024)

Also indexed in: Semantic Scholar · OpenAlex · dblp · Publisher page

TL;DR. S3CFSL classifies new hyperspectral scenes from five labelled samples per class by transferring knowledge from the labelled Chikusei scene. It combines a cross-spatial–spectral transformer, Gaussian feature denoising and semantic-enhanced domain alignment, and reports 98.52% overall accuracy on Pavia Centre, 88.79% on Salinas and 77.63% on Houston.

Quick facts

Method nameS3CFSL; S³CFSL (typeset form with superscript 3 used in the published IEEE PDF)
AuthorsMengxin Cao, Xu Zhang, Jinyong Cheng, Guixin Zhao, Wei Li, Xiangjun Dong
DOI10.1109/TGRS.2024.3434484
TaskCross-domain few-shot hyperspectral image (HSI) classification, where source and target scenes have non-overlapping class label sets
Problems addressedResidual noise left after preprocessing, and sample-selection bias that can create false statistical correlations between domains and reduce generalisation
Feature extractorSpatial and spectral dual channels (SSDC) feeding a cross-spatial–spectral transformer (CSST) built from spatial–spectral cross-attention (SSC attention) and a spatial–spectral fusion extractor (SSFE)
Feature denoisingA Gaussian filter is applied to the extracted features and its output is added back to the unfiltered features to avoid over-blurring
Domain alignmentSemantic-enhanced domain alignment (SEDA): convolutional semantic masks and k-means clustering (k = 2) on top of a conditional domain adversarial network
DatasetsSource: Chikusei (19 classes, 128 bands). Targets: Pavia Centre (9 classes), Salinas (16 classes), Houston (15 classes)
Protocol and metricsFive labelled target samples per class; 1-shot, C-way training episodes; 9 × 9 input patches; overall accuracy (OA), average accuracy (AA) and Kappa
BaselinesSS CNN, DFSL, DCFSL, CMFSL, Gia-FSL, GCC-FSL, DMCM and SSTFSL (the ICIP 2023 conference version of this work)
Headline resultOA 98.52% (Pavia Centre), 88.79% (Salinas), 77.63% (Houston); higher OA, AA and Kappa than all eight baselines (Tables V–VII)
Corresponding authorsGuixin Zhao and Xiangjun Dong
VenueIEEE Transactions on Geoscience and Remote Sensing, vol. 62, 2024, article no. 5525315 (15 pages); published online 29 July 2024; IEEE Xplore document 10613782
AccessNot open access: subscription article in IEEE Xplore, with no arXiv or repository copy found
BibTeX keycao2024s3fsl

Main results

SettingResultCompared with
Pavia Centre target (source Chikusei), 5 labelled samples per class, Table VOA = 98.52%, AA = 95.77%, Kappa = 0.9728
“Different Classification Methods on the PC Dataset”
GCC-FSL (strongest baseline): OA 98.45%, AA 95.53%, Kappa 0.9716; margin 0.07 OA points
Salinas target (source Chikusei), 5 labelled samples per class, Table VIOA = 88.79%, AA = 90.41%, Kappa = 0.8725
“Different Classification Methods on the SA Dataset”
GCC-FSL: OA 88.33%, AA 90.04%, Kappa 0.8711; highest baseline AA is Gia-FSL at 90.11%
Houston target (source Chikusei), 5 labelled samples per class, Table VIIOA = 77.63%, AA = 79.93%, Kappa = 0.7583
“Different Classification Methods on the Houston Dataset”
GCC-FSL: OA 77.44%, AA 79.76%, Kappa 0.7536; margin 0.19 OA points
Module ablation, OA on Pavia Centre / Salinas / Houston, Table VIIIFull S3CFSL OA = 98.52 / 88.79 / 77.63
“Ablation Analysis of the S3CFSL With a Combination of Different Modules for the Three Datasets on OA”
Plain ViT extractor without denoising or SEDA: 94.33 / 81.46 / 69.77; SSDC + CSST with denoising but without SEDA: 97.25 / 87.44 / 75.69
Denoising-filter ablation, OA on Pavia Centre / Salinas / Houston, Table IXGaussian filter OA = 98.52 / 88.79 / 77.63
“Gaussian filtering is the best among them.”
Median filter 98.43 / 88.74 / 77.51; mean filter 98.48 / 88.55 / 77.08; no filter 98.36 / 88.06 / 77.14 (mean filter falls below no filter on Houston)

Quoted text is verbatim from the paper.

Key points

Abstract

Preprocessing procedures are commonly employed to reduce water-absorption bands and noise in hyperspectral images (HSIs). Nevertheless, they typically do not entirely eradicate noise. This is especially evident in scenarios that necessitate data of exceptional quality, such as cross-domain few-shot classification tasks. Within these specific conditions, the influence of remaining background noise on the ultimate results of classification is substantial. Furthermore, the presence of sample selection biases in the few-shot task might lead to the emergence of false statistical correlations between data from distinct domains, resulting in a decrease in the model’s ability to generalize. We propose a new method called spatial-spectral–semantic cross-domain few-shot learning (S3CFSL) to address the challenge. This method promotes the learning of transferable information by incorporating feature denoising operations in the feature extraction process to restore essential information. Concurrently, it enhances cross-domain distributional consistency by introducing a semantic-aware strategy to strengthen the association between cross-domain data and semantic information. Specifically, the spatial and spectral dual channels (SSDCs), in conjunction with the cross-spatial-spectral transformer (CSST), are designed as a feature extractor to acquire interactive spatial-spectral features. The feature-denoising operations can further acquiring transferable information from cross-domain features, thus facilitating meta-learning in both the source domain (SD) and the target domain (TD). Meanwhile, a semantic-enhanced domain alignment (SEDA) is designed to promote domain adaptation by using a semantic-aware strategy, which significantly enhances distributional consistency for cross-domain tasks. Our results exhibit exceptional classification efficacy in comparison to other state-of-the-art approaches on three public HSI datasets.

Frequently asked questions

What is cross-domain few-shot hyperspectral image classification?

Classifying a new hyperspectral scene (the target domain) that has only a few labelled samples, using knowledge transferred from a different scene with plentiful labels (the source domain) whose class label set does not overlap with the target's. S3CFSL alternates few-shot episodes on the source and target domains.

How does S3CFSL improve cross-domain transfer?

Feature-denoising operations restore essential information for learning transferable features. They apply a Gaussian filter and add its output back to the unfiltered features. Semantic-enhanced domain alignment (SEDA) uses a semantic-aware strategy (semantic masks and k-means clustering, built on a conditional domain adversarial network) to strengthen the link between cross-domain data and semantic information, which improves cross-domain distributional consistency.

Which datasets and few-shot setting does S3CFSL use?

Chikusei (19 classes, 128 bands) is the labelled source domain. The three target domains are Pavia Centre (ROSIS sensor, 9 classes, 102 bands), Salinas (AVIRIS sensor, 16 classes, 204 bands after 20 water-absorption bands are removed) and Houston (CASI sensor, University of Houston campus, 15 classes, 144 bands). The main comparison uses five labelled target samples per class, with 1-shot, C-way training episodes where C is the number of target classes.

What accuracy does S3CFSL achieve?

With five labelled samples per target class, S3CFSL reports overall accuracy (OA) of 98.52% on Pavia Centre, 88.79% on Salinas and 77.63% on Houston, with average accuracy of 95.77%, 90.41% and 79.93% and Kappa of 0.9728, 0.8725 and 0.7583. These OA, AA and Kappa values are higher than those of all eight compared methods (SS CNN, DFSL, DCFSL, CMFSL, Gia-FSL, GCC-FSL, DMCM and SSTFSL) on every target dataset. The closest competitor, GCC-FSL, reaches 98.45%, 88.33% and 77.44% OA, so the OA margins are 0.07 to 0.46 points.

What does the S3CFSL ablation show?

With a plain ViT feature extractor and neither feature denoising nor SEDA, OA is 94.33 / 81.46 / 69.77 on Pavia Centre / Salinas / Houston. Switching to the SSDC + CSST extractor gives 96.47 / 86.39 / 75.16, and adding feature denoising gives 97.25 / 87.44 / 75.69. The full S3CFSL with SEDA reaches 98.52 / 88.79 / 77.63, and in the filter ablation the Gaussian filter beats median, mean and no filtering on all three datasets.

How is S3CFSL related to SSTFSL?

SSTFSL (Cao et al., ICIP 2023, "Few-Shot Hyperspectral Image Classification Based on Cross-Domain Spectral Semantic Relation Transformer", DOI 10.1109/ICIP49359.2023.10222564) is the preliminary conference version and a separate publication with a partly different author list. S3CFSL adds the CSST with SSDC feature extractor, feature denoising, a semantic-aware strategy for domain alignment, and validation on additional public datasets. Its OA is higher than that of SSTFSL on Pavia Centre (98.52 vs 97.85), Salinas (88.79 vs 86.86) and Houston (77.63 vs 76.82).

How do I cite S3CFSL?

Cite it as: Cao, M., Zhang, X., Cheng, J., Zhao, G., Li, W., & Dong, X. (2024). Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensing, 62, Article 5525315. https://doi.org/10.1109/TGRS.2024.3434484 BibTeX is available at https://codezx6.github.io/papers/spatial-spectral-semantic-fsl.html#cite and in https://codezx6.github.io/publications.bib.

中文摘要

面向高光谱图像分类的空间-光谱-语义跨域小样本学习(S3CFSL)

S3CFSL 用于跨域小样本高光谱图像分类:以空间与光谱双通道和跨空间-光谱 Transformer 提取特征,高斯滤波去噪后与原特征相加,并以语义增强域对齐提升分布一致性。以 Chikusei 为源域、每类 5 个标注样本,在 Pavia Centre、Salinas、Houston 上总体精度为 98.52%、88.79%、77.63%,高于八种对比方法。

关键词:高光谱图像分类;跨域小样本学习;域适应;分布对齐;特征去噪;语义增强域对齐

Keywords

hyperspectral image classification few-shot learning cross-domain domain adaptation distribution alignment spatial-spectral semantic-enhanced domain alignment remote sensing

Also referred to as: S³CFSL; spatial-spectral-semantic cross-domain few-shot learning; Spatial–Spectral–Semantic Cross-Domain FSL for HSI Classification; spatial and spectral dual channels (SSDC); cross-spatial–spectral transformer (CSST); semantic-enhanced domain alignment (SEDA).

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › computer vision › remote sensing › Earth observation › hyperspectral imaging › hyperspectral image classification › few-shot hyperspectral classification › cross-domain few-shot HSI classification › S3CFSL

Task

hyperspectral image classification HSI classification cross-domain few-shot learning few-shot hyperspectral classification domain adaptation for remote sensing remote sensing image classification

Method

meta-learning (episodic training) feature denoising semantic-aware domain alignment spatial-spectral Transformer spatial-spectral feature extraction

Datasets

Chikusei Pavia Centre Salinas Houston

Compared with

DFSL DCFSL CMFSL Gia-FSL GCC-FSL DMCM

中文

高光谱图像分类 高光谱遥感 跨域小样本学习 小样本学习 域适应 元学习 遥感图像分类

Related areas

few-shot learning meta-learning transfer learning domain adaptation domain generalization label-efficient learning Vision Transformers spectral-spatial feature learning denoising land cover classification multispectral imaging

Applications

land-cover mapping precision agriculture mineral exploration environmental monitoring urban mapping

Related benchmarks (not used in this paper)

Indian Pines Pavia University Kennedy Space Center (KSC) Botswana WHU-Hi

中文 (扩展)

遥感 对地观测 高光谱成像 小样本高光谱分类 迁移学习 域泛化 标签高效学习 视觉Transformer 空谱特征学习 去噪 地物分类 土地覆盖制图 环境监测

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Hyperspectral image classification

Other topics

Cite this paper

@article{cao2024s3fsl,
  title        = {Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification},
  author       = {Cao, Mengxin and Zhang, Xu and Cheng, Jinyong and Zhao, Guixin and Li, Wei and Dong, Xiangjun},
  journal      = {IEEE Transactions on Geoscience and Remote Sensing},
  year         = {2024},
  volume       = {62},
  pages        = {5525315},
  doi          = {10.1109/TGRS.2024.3434484},
  issn         = {0196-2892},
  url          = {https://doi.org/10.1109/TGRS.2024.3434484}
}
Cao, M., Zhang, X., Cheng, J., Zhao, G., Li, W., & Dong, X. (2024). Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensing, 62, Article 5525315. https://doi.org/10.1109/TGRS.2024.3434484