Papers on speech emotion recognition, emotional TTS, urban flow and OD demand prediction, and hyperspectral image classification
Each paper page gives a plain-language summary, quick facts, main results with verbatim quotes, the abstract, an FAQ, a Chinese summary, and copyable BibTeX and APA citations. Research area map · All BibTeX · 中文版 · Atom feed
Speech emotion recognition and emotional text-to-speech for humanoid robots
Two 2026 arXiv papers address both directions of affective speech: recognising emotion in speech (PhysioSER; Zhang et al., 2026) and generating emotional speech (DUET; Zhang et al., 2026). Both methods are designed as plug-and-play frameworks that keep the pretrained model frozen, and both were deployed on a humanoid robot.
DUET adds emotion control to frozen diffusion and flow-matching TTS models by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients passed through a differentiable vocoder. With GradTTS on ESD it reaches 75.5% average emotion accuracy (strongest supervised baseline: 46.8%).
PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.
Urban flow prediction with masked and contrastive spatio-temporal pre-training
Urban flow prediction forecasts future flows across city regions from historical flow data and supports services such as trip planning, congestion control, and public safety. Three papers (Zhang et al., CIKM 2023; Zhang et al., Knowledge-Based Systems 2023; Zhang et al., Neural Networks 2025) develop self-supervised spatio-temporal pre-training for this task.
MR-UFP predicts the inflow and outflow of city grid regions, even with limited training data, by pre-training encoders with spatial-temporal masking and contrastive learning and training flow prediction jointly with a multi-scale region-classification task. On full TaxiBJ and BikeNYC it beats all 13 baselines (RMSE 14.32 and 4.45).
ST-FCL predicts grid-level urban inflow and outflow by fusing temporal and spatial views, learned through contrastive pretraining, with an external-factor view. On the full TaxiBJ dataset it reaches RMSE 14.71, against 15.41 for the best baseline, ATFM.
MC-STL pre-trains two encoders for urban flow prediction: a ViT encoder learns to reconstruct regions masked at different timestamps, and its attention weights also build a GCN adjacency matrix; a global-local cross-attention encoder learns a temporal-order contrastive task. It reaches RMSE 14.53 on full TaxiBJ.
Origin-destination (OD) flow and demand prediction for metro and taxi systems
Origin-destination (OD) prediction forecasts the flow or demand between pairs of stations or regions, which supports real-time traffic management, vehicle dispatch, and resource allocation in intelligent transportation systems. Two papers (Yu et al., Expert Systems with Applications 2025; Gong et al., Expert Systems with Applications 2026) study this problem on metro and taxi data.
DSTCN (Dynamic Spatio-Temporal Correlation Network) forecasts origin-destination demand matrices with three modules: Glstm2D for origin- and destination-side demand trends, Simformer for Transformer-based spatial similarity across the OD matrix, and FF-TM for feature fusion and temporal modeling. On NYC-TOD2018 its MAE of 1.469 is 3.04% below the best baseline.
BiST-IF predicts origin–destination (OD) flows between metro stations or urban areas by correcting delayed recent OD matrices, applying bi-directional origin/destination attention, and fusing arrival-side (Out-OD) flows through an attention-based mutual information mechanism. It lowers MAE by an average of 7.55% on HZMetro relative to the best baseline.
Hyperspectral image classification: cross-domain few-shot learning and superpixel mutual enhancement
Two IEEE Transactions on Geoscience and Remote Sensing papers (Cao et al., 2024; Cao et al., 2025) study hyperspectral image (HSI) classification.
S²CMEN (S2CMEN) is a hyperspectral image classification network in which superpixel segmentation and classification enhance each other through a unified loss, combining superpixel-based global spatial context with Spectral-Swin Transformer spectral features. It reaches 97.38% overall accuracy on Indian Pines.
S3CFSL classifies new hyperspectral scenes from five labelled samples per class by transferring knowledge from the labelled Chikusei scene. It combines a cross-spatial–spectral transformer, Gaussian feature denoising and semantic-enhanced domain alignment, and reports 98.52% overall accuracy on Pavia Centre, 88.79% on Salinas and 77.63% on Houston.
Object detection
A tomato leaf disease detector that adds the DyHead attention module to YOLOv4-tiny and trains it with a Focaler-SIoU box loss. On PlantDoc tomato leaf images it reaches 93.64% mAP, 10.3 percentage points above the YOLOv4-tiny baseline.
URMDet-SimFire is an uncertainty- and reliability-aware multimodal object detection framework for fire and smoke that reweights each input modality by its estimated reliability and calibrates each detection's confidence by its uncertainty. In simulation only, it raised F1 from 0.4348 (vision-only baseline) to 0.4638.