Papers on speech emotion recognition, emotional TTS, urban flow and OD demand prediction, and hyperspectral image classification

Each paper page gives a plain-language summary, quick facts, main results with verbatim quotes, the abstract, an FAQ, a Chinese summary, and copyable BibTeX and APA citations. Research area map · All BibTeX · 中文版 · Atom feed

Speech emotion recognition and emotional text-to-speech for humanoid robots

Two 2026 arXiv papers address both directions of affective speech: recognising emotion in speech (PhysioSER; Zhang et al., 2026) and generating emotional speech (DUET; Zhang et al., 2026). Both methods are designed as plug-and-play frameworks that keep the pretrained model frozen, and both were deployed on a humanoid robot.

Xu Zhang, Longbing Cao, Zhangkai Wu
arXiv:2606.00066 (2026)

DUET adds emotion control to frozen diffusion and flow-matching TTS models by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients passed through a differentiable vocoder. With GradTTS on ESD it reaches 75.5% average emotion accuracy (strongest supervised baseline: 46.8%).

Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu
arXiv:2602.13259 (2026)

PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.

Urban flow prediction with masked and contrastive spatio-temporal pre-training

Urban flow prediction forecasts future flows across city regions from historical flow data and supports services such as trip planning, congestion control, and public safety. Three papers (Zhang et al., CIKM 2023; Zhang et al., Knowledge-Based Systems 2023; Zhang et al., Neural Networks 2025) develop self-supervised spatio-temporal pre-training for this task.

Xu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang
Neural Networks, vol. 182, article 106900 (2025)

MR-UFP predicts the inflow and outflow of city grid regions, even with limited training data, by pre-training encoders with spatial-temporal masking and contrastive learning and training flow prediction jointly with a multi-scale region-classification task. On full TaxiBJ and BikeNYC it beats all 13 baselines (RMSE 14.32 and 4.45).

Xu Zhang, Yongshun Gong, Chengqi Zhang, Xiaoming Wu, Ying Guo, Wenpeng Lu, Long Zhao, Xiangjun Dong
Knowledge-Based Systems, vol. 282, article 111104 (2023)

ST-FCL predicts grid-level urban inflow and outflow by fusing temporal and spatial views, learned through contrastive pretraining, with an external-factor view. On the full TaxiBJ dataset it reaches RMSE 14.71, against 15.41 for the best baseline, ATFM.

Xu Zhang, Yongshun Gong, Xinxin Zhang, Xiaoming Wu, Chengqi Zhang, Xiangjun Dong
CIKM 2023, pp. 3298–3307

MC-STL pre-trains two encoders for urban flow prediction: a ViT encoder learns to reconstruct regions masked at different timestamps, and its attention weights also build a GCN adjacency matrix; a global-local cross-attention encoder learns a temporal-order contrastive task. It reaches RMSE 14.53 on full TaxiBJ.

Origin-destination (OD) flow and demand prediction for metro and taxi systems

Origin-destination (OD) prediction forecasts the flow or demand between pairs of stations or regions, which supports real-time traffic management, vehicle dispatch, and resource allocation in intelligent transportation systems. Two papers (Yu et al., Expert Systems with Applications 2025; Gong et al., Expert Systems with Applications 2026) study this problem on metro and taxi data.

Yongshun Gong, Piao Yu, Xu Zhang, Xinxin Zhang, Xiushan Nie, Haoliang Sun
Expert Systems with Applications, vol. 299, article 130095 (2026)

DSTCN (Dynamic Spatio-Temporal Correlation Network) forecasts origin-destination demand matrices with three modules: Glstm2D for origin- and destination-side demand trends, Simformer for Transformer-based spatial similarity across the OD matrix, and FF-TM for feature fusion and temporal modeling. On NYC-TOD2018 its MAE of 1.469 is 3.04% below the best baseline.

Piao Yu, Xu Zhang, Yongshun Gong, Jian Zhang, Haoliang Sun, Junjie Zhang, Xinxin Zhang, Yilong Yin
Expert Systems with Applications, vol. 264, article 125679 (2025)

BiST-IF predicts origin–destination (OD) flows between metro stations or urban areas by correcting delayed recent OD matrices, applying bi-directional origin/destination attention, and fusing arrival-side (Out-OD) flows through an attention-based mutual information mechanism. It lowers MAE by an average of 7.55% on HZMetro relative to the best baseline.

Hyperspectral image classification: cross-domain few-shot learning and superpixel mutual enhancement

Two IEEE Transactions on Geoscience and Remote Sensing papers (Cao et al., 2024; Cao et al., 2025) study hyperspectral image (HSI) classification.

Mengxin Cao, Yongmin Li, Xu Zhang, Guixin Zhao, Guohua Lv, Aimei Dong, Jinyong Cheng, Wei Li, Xiangjun Dong
IEEE Transactions on Geoscience and Remote Sensing, vol. 63, article 5524714 (2025)

S²CMEN (S2CMEN) is a hyperspectral image classification network in which superpixel segmentation and classification enhance each other through a unified loss, combining superpixel-based global spatial context with Spectral-Swin Transformer spectral features. It reaches 97.38% overall accuracy on Indian Pines.

Mengxin Cao, Xu Zhang, Jinyong Cheng, Guixin Zhao, Wei Li, Xiangjun Dong
IEEE Transactions on Geoscience and Remote Sensing, vol. 62, article 5525315 (2024)

S3CFSL classifies new hyperspectral scenes from five labelled samples per class by transferring knowledge from the labelled Chikusei scene. It combines a cross-spatial–spectral transformer, Gaussian feature denoising and semantic-enhanced domain alignment, and reports 98.52% overall accuracy on Pavia Centre, 88.79% on Salinas and 77.63% on Houston.

Object detection

Yumeng Yao, Xiaodun Deng, Xu Zhang, Junming Li, Wenxuan Sun, Gechao Zhang
PeerJ Computer Science, vol. 10, article e2365 (2024)

A tomato leaf disease detector that adds the DyHead attention module to YOLOv4-tiny and trains it with a Focaler-SIoU box loss. On PlantDoc tomato leaf images it reaches 93.64% mAP, 10.3 percentage points above the YOLOv4-tiny baseline.

Zhonghua Dang, Xu Zhang
Array, vol. 31, article 101124 (2026)

URMDet-SimFire is an uncertainty- and reliability-aware multimodal object detection framework for fire and smoke that reweights each input modality by its estimated reliability and calibrates each detection's confidence by its uncertainty. In simulation only, it raised F1 from 0.4348 (vision-only baseline) to 0.4638.