<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>Research paper pages</title>
<link href="https://codezx6.github.io/feed.xml" rel="self"/>
<link href="https://pubsubhubbub.appspot.com/" rel="hub"/>
<link href="https://codezx6.github.io/"/>
<id>https://codezx6.github.io/</id>
<updated>2026-09-27T00:00:00Z</updated>
<author><name>Research paper pages</name></author>
<entry><title>URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics</title><link href="https://codezx6.github.io/papers/urmdet-simfire.html"/><id>https://codezx6.github.io/papers/urmdet-simfire.html</id><published>2026-08-07T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Zhonghua Dang</name></author><author><name>Xu Zhang</name></author><summary>URMDet-SimFire is an uncertainty- and reliability-aware multimodal object detection framework for fire and smoke that reweights each input modality by its estimated reliability and calibrates each detection&#x27;s confidence by its uncertainty. In simulation only, it raised F1 from 0.4348 (vision-only baseline) to 0.4638.</summary><category term="trustworthy AI"/><category term="multimodal fusion"/><category term="uncertainty quantification"/><category term="reliability learning"/><category term="safety-critical perception"/><category term="object detection"/></entry>
<entry><title>DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech</title><link href="https://codezx6.github.io/papers/duet.html"/><id>https://codezx6.github.io/papers/duet.html</id><published>2026-05-20T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Xu Zhang</name></author><author><name>Longbing Cao</name></author><author><name>Zhangkai Wu</name></author><summary>DUET adds emotion control to frozen diffusion and flow-matching TTS models by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients passed through a differentiable vocoder. With GradTTS on ESD it reaches 75.5% average emotion accuracy (strongest supervised baseline: 46.8%).</summary><category term="emotional text-to-speech"/><category term="emotion control"/><category term="diffusion-based TTS"/><category term="flow-matching TTS"/><category term="representation steering"/><category term="classifier guidance"/></entry>
<entry><title>PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition</title><link href="https://codezx6.github.io/papers/physioser.html"/><id>https://codezx6.github.io/papers/physioser.html</id><published>2026-02-03T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Xu Zhang</name></author><author><name>Longbing Cao</name></author><author><name>Runze Yang</name></author><author><name>Zhangkai Wu</name></author><summary>PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.</summary><category term="speech emotion recognition"/><category term="physiology-informed vocal representation"/><category term="quaternion neural networks"/><category term="self-supervised learning"/><category term="amplitude and phase"/><category term="group delay and instantaneous frequency"/></entry>
<entry><title>DSTCN: Exploiting dynamic spatio-temporal correlations for origin-destination demand prediction</title><link href="https://codezx6.github.io/papers/dstcn.html"/><id>https://codezx6.github.io/papers/dstcn.html</id><published>2025-10-24T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Yongshun Gong</name></author><author><name>Piao Yu</name></author><author><name>Xu Zhang</name></author><author><name>Xinxin Zhang</name></author><author><name>Xiushan Nie</name></author><author><name>Haoliang Sun</name></author><summary>DSTCN (Dynamic Spatio-Temporal Correlation Network) forecasts origin-destination demand matrices with three modules: Glstm2D for origin- and destination-side demand trends, Simformer for Transformer-based spatial similarity across the OD matrix, and FF-TM for feature fusion and temporal modeling. On NYC-TOD2018 its MAE of 1.469 is 3.04% below the best baseline.</summary><category term="origin-destination demand prediction"/><category term="spatio-temporal modeling"/><category term="Transformer"/><category term="graph neural network"/><category term="intelligent transportation systems"/><category term="OD demand matrix"/></entry>
<entry><title>S2CMEN: A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image</title><link href="https://codezx6.github.io/papers/mutual-enhancement-hsi.html"/><id>https://codezx6.github.io/papers/mutual-enhancement-hsi.html</id><published>2025-09-11T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Mengxin Cao</name></author><author><name>Yongmin Li</name></author><author><name>Xu Zhang</name></author><author><name>Guixin Zhao</name></author><author><name>Guohua Lv</name></author><author><name>Aimei Dong</name></author><author><name>Jinyong Cheng</name></author><author><name>Wei Li</name></author><author><name>Xiangjun Dong</name></author><summary>S²CMEN (S2CMEN) is a hyperspectral image classification network in which superpixel segmentation and classification enhance each other through a unified loss, combining superpixel-based global spatial context with Spectral-Swin Transformer spectral features. It reaches 97.38% overall accuracy on Indian Pines.</summary><category term="hyperspectral image classification"/><category term="superpixel segmentation"/><category term="feature fusion"/><category term="Transformer"/><category term="graph convolutional network"/><category term="mutual enhancement"/></entry>
<entry><title>BiST-IF: Enhancing origin–destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution</title><link href="https://codezx6.github.io/papers/bist-if.html"/><id>https://codezx6.github.io/papers/bist-if.html</id><published>2024-11-22T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Piao Yu</name></author><author><name>Xu Zhang</name></author><author><name>Yongshun Gong</name></author><author><name>Jian Zhang</name></author><author><name>Haoliang Sun</name></author><author><name>Junjie Zhang</name></author><author><name>Xinxin Zhang</name></author><author><name>Yilong Yin</name></author><summary>BiST-IF predicts origin–destination (OD) flows between metro stations or urban areas by correcting delayed recent OD matrices, applying bi-directional origin/destination attention, and fusing arrival-side (Out-OD) flows through an attention-based mutual information mechanism. It lowers MAE by an average of 7.55% on HZMetro relative to the best baseline.</summary><category term="Origin–destination flow prediction"/><category term="Spatio-temporal data"/><category term="Bi-directional attention mechanism"/><category term="Traffic prediction"/><category term="Intelligent transport systems"/><category term="OD delay correction"/></entry>
<entry><title>MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information</title><link href="https://codezx6.github.io/papers/mr-ufp.html"/><id>https://codezx6.github.io/papers/mr-ufp.html</id><published>2024-11-16T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Xu Zhang</name></author><author><name>Mengxin Cao</name></author><author><name>Yongshun Gong</name></author><author><name>Xiaoming Wu</name></author><author><name>Xiangjun Dong</name></author><author><name>Ying Guo</name></author><author><name>Long Zhao</name></author><author><name>Chengqi Zhang</name></author><summary>MR-UFP predicts the inflow and outflow of city grid regions, even with limited training data, by pre-training encoders with spatial-temporal masking and contrastive learning and training flow prediction jointly with a multi-scale region-classification task. On full TaxiBJ and BikeNYC it beats all 13 baselines (RMSE 14.32 and 4.45).</summary><category term="urban flow prediction"/><category term="mutual reinforcement"/><category term="multi-scale region information"/><category term="spatial–temporal systems"/><category term="spatial-temporal random masking"/><category term="contrastive pre-training"/></entry>
<entry><title>Automatic visual recognition for leaf disease based on enhanced attention mechanism</title><link href="https://codezx6.github.io/papers/leaf-disease-attention.html"/><id>https://codezx6.github.io/papers/leaf-disease-attention.html</id><published>2024-11-04T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Yumeng Yao</name></author><author><name>Xiaodun Deng</name></author><author><name>Xu Zhang</name></author><author><name>Junming Li</name></author><author><name>Wenxuan Sun</name></author><author><name>Gechao Zhang</name></author><summary>A tomato leaf disease detector that adds the DyHead attention module to YOLOv4-tiny and trains it with a Focaler-SIoU box loss. On PlantDoc tomato leaf images it reaches 93.64% mAP, 10.3 percentage points above the YOLOv4-tiny baseline.</summary><category term="visual recognition"/><category term="leaf disease identification"/><category term="attention mechanism"/><category term="tomato leaf disease"/><category term="object detection"/><category term="YOLOv4-tiny"/></entry>
<entry><title>S3CFSL: Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification</title><link href="https://codezx6.github.io/papers/spatial-spectral-semantic-fsl.html"/><id>https://codezx6.github.io/papers/spatial-spectral-semantic-fsl.html</id><published>2024-07-29T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Mengxin Cao</name></author><author><name>Xu Zhang</name></author><author><name>Jinyong Cheng</name></author><author><name>Guixin Zhao</name></author><author><name>Wei Li</name></author><author><name>Xiangjun Dong</name></author><summary>S3CFSL classifies new hyperspectral scenes from five labelled samples per class by transferring knowledge from the labelled Chikusei scene. It combines a cross-spatial–spectral transformer, Gaussian feature denoising and semantic-enhanced domain alignment, and reports 98.52% overall accuracy on Pavia Centre, 88.79% on Salinas and 77.63% on Houston.</summary><category term="hyperspectral image classification"/><category term="few-shot learning"/><category term="cross-domain"/><category term="domain adaptation"/><category term="distribution alignment"/><category term="spatial-spectral"/></entry>
<entry><title>ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction</title><link href="https://codezx6.github.io/papers/st-csl.html"/><id>https://codezx6.github.io/papers/st-csl.html</id><published>2023-10-21T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Xu Zhang</name></author><author><name>Yongshun Gong</name></author><author><name>Chengqi Zhang</name></author><author><name>Xiaoming Wu</name></author><author><name>Ying Guo</name></author><author><name>Wenpeng Lu</name></author><author><name>Long Zhao</name></author><author><name>Xiangjun Dong</name></author><summary>ST-FCL predicts grid-level urban inflow and outflow by fusing temporal and spatial views, learned through contrastive pretraining, with an external-factor view. On the full TaxiBJ dataset it reaches RMSE 14.71, against 15.41 for the best baseline, ATFM.</summary><category term="urban flow prediction"/><category term="crowd flow prediction"/><category term="contrastive learning"/><category term="spatio-temporal fusion"/><category term="multi-view representation fusion"/><category term="multi-view influencing factors"/></entry>
<entry><title>MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction</title><link href="https://codezx6.github.io/papers/mcstl.html"/><id>https://codezx6.github.io/papers/mcstl.html</id><published>2023-10-21T00:00:00Z</published><updated>2026-09-27T00:00:00Z</updated><author><name>Xu Zhang</name></author><author><name>Yongshun Gong</name></author><author><name>Xinxin Zhang</name></author><author><name>Xiaoming Wu</name></author><author><name>Chengqi Zhang</name></author><author><name>Xiangjun Dong</name></author><summary>MC-STL pre-trains two encoders for urban flow prediction: a ViT encoder learns to reconstruct regions masked at different timestamps, and its attention weights also build a GCN adjacency matrix; a global-local cross-attention encoder learns a temporal-order contrastive task. It reaches RMSE 14.53 on full TaxiBJ.</summary><category term="urban flow prediction"/><category term="spatio-temporal predictive modeling"/><category term="spatio-temporal pre-training"/><category term="mask-enhanced learning"/><category term="contrastive learning"/><category term="traffic prediction"/></entry>
</feed>
