MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information
Xu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang
Neural Networks, vol. 182, article 106900 (2025)
Also indexed in: Semantic Scholar · OpenAlex · dblp · PubMed · Publisher page · PolyU Institutional Research Archive
Quick facts
| Method name | MR-UFP; MR-UPF (GitHub repository name, github.com/CodeZx6/MR-UPF; the paper's data footnote links github.com/CodeZx6/Data-of-MR-UPF, which now redirects there) |
|---|---|
| Authors | Xu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang |
| DOI | 10.1016/j.neunet.2024.106900 |
| Task | Urban flow prediction: forecasting the inflow and outflow of each grid region in a city, a key component of intelligent transportation systems |
| Pre-training | EncM, a ViT encoder pre-trained to reconstruct spatio-temporally masked regions (75% masking rate), and EncC, a global–local cross-attention encoder pre-trained with a triplet loss on temporally shuffled negatives |
| Auxiliary task | Multi-scale region classification (road/non-road, coarse, fine-grained) on hand-labeled region texts, trained jointly with flow prediction so the two tasks reinforce each other |
| Other modules | Time-text and Region-text Encoders, adaptive environment feature fusion, and a real-time graph processing module (RGPM) that turns EncM attention weights into an adjacency matrix for a GCN |
| Joint loss | Three region losses combined with learnable weights into L_rcf; Loss = 2·L_P·L_rcf/(L_P + L_rcf), so the smaller loss gets more weight and neither dominates |
| Datasets | TaxiBJ (Beijing taxis, 2×32×32 grid, 30-min intervals, holiday and weather data) and BikeNYC (New York City bikes, 2×16×8 grid, 1-h intervals, no weather data) |
| Region labels | Hand-labeled from real maps: 2/5/142 classes (TaxiBJ) and 2/8/24 (BikeNYC) at the three scales; released with the code at github.com/CodeZx6/MR-UPF |
| Metrics and baselines | RMSE and MAE averaged over 5 random seeds, with t-tests; 13 baselines including ST-ResNet, ATFM, AGCRN, STGCN, PDFormer, STAEformer, ST-FCL, MC-STL and DeepMeshCity |
| Headline result | Full TaxiBJ RMSE/MAE 14.32/8.56 vs 14.53/8.70 for MC-STL; full BikeNYC 4.45/2.19 vs 4.63/2.21 for DeepMeshCity |
| Authorship | Xu Zhang and Mengxin Cao contributed equally; Xiaoming Wu and Xiangjun Dong are corresponding authors |
| Publication | Neural Networks, vol. 182, article 106900 (February 2025 issue); received 19 June 2024, accepted 7 November 2024, available online 16 November 2024 |
| Access | Subscription article, not open access. The PolyU Institutional Research Archive lists an accepted manuscript under embargo until 28 February 2027 |
| Code | https://github.com/CodeZx6/MR-UPF |
| BibTeX key | zhang2025mrufp |
Main results
| Setting | Result | Compared with |
|---|---|---|
| TaxiBJ, full dataset (last 28 days as test set) | RMSE = 14.32, MAE = 8.56“our proposed MR-UFP model yielded a mean RMSE of 14.32 and an MAE of 8.56. This represents an enhancement in prediction accuracy by 1.47% and 1.64% over the sub-optimal model MC-STL, respectively.” | MC-STL (strongest baseline): RMSE 14.53, MAE 8.70; p < 0.01 |
| BikeNYC, full dataset (last 10 days as test set; no weather data) | RMSE = 4.45, MAE = 2.19“In the experiment utilizing the complete BikeNYC dataset, we used the last 10 days for testing. As shown in Table 2, MR-UFP achieved an RMSE of 4.45 and an MAE of 2.19.” | DeepMeshCity (strongest baseline): RMSE 4.63, MAE 2.21; Table 2 gives p < 0.01 (RMSE) and p < 0.03 (MAE) |
| BikeNYC, 20% data subset (last 10% as test set) | RMSE = 4.15, MAE = 2.09“MR-UFP (Ours) 4.15 ± 0.03 2.09 ± 0.01” | DeepMeshCity: RMSE 4.55, MAE 2.13 |
| TaxiBJ P3, 20% data subset (last 10% as test set) | RMSE = 15.57, MAE = 9.95“MR-UFP(Ours) 15.57 ± 0.04 9.95 ± 0.03” | Best baselines: ST-FCL RMSE 16.82; MC-STL MAE 10.77 |
| TaxiBJ subsets P3(80%), P3(100%) and P4(20%) | RMSE rank = second-best; MAE rank = best“Notably, in the datasets P3(80%), P3(100%), and P4(20%), although the RMSE of MR-UFP was sub-optimal, the MAE reached optimal levels.” | — |
| Ablation, full TaxiBJ P1 | Full model RMSE = 15.36 (MAE 9.16); without region classification (w/o-MultiFC) RMSE = 15.80 (MAE 9.35)“the RMSE loss of the variant ‘MR-UFP-w/o-MultiFC’ increases to 15.80 when this auxiliary task is entirely removed” | Without EncM branch (w/o-STMask) 15.80; without EncC branch (w/o-STCon) 15.72; without environment features 15.66; without both pre-trainings 15.52; without masked pre-training only 15.45 (RMSE) |
Quoted text is verbatim from the paper.
Key points
- Two encoders are pre-trained on historical flows: EncM, a ViT encoder that learns to reconstruct spatio-temporally masked regions (75% masking worked best), and EncC, a global–local cross-attention encoder trained with a triplet loss.
- Regions are hand-labeled from real maps at three scales (road/non-road, coarse, fine-grained; 2/5/142 classes on TaxiBJ), and classifying them is an auxiliary task trained jointly with flow prediction.
- A joint loss, 2·L_P·L_rcf/(L_P + L_rcf), keeps either loss from dominating; on TaxiBJ P1 it gives RMSE 15.36 versus 15.41 for a plain L_P + L_rcf sum.
- On full TaxiBJ and BikeNYC, MR-UFP beats all 13 baselines (RMSE 14.32 vs 14.53 for MC-STL; 4.45 vs 4.63 for DeepMeshCity) and is best on most 20–100% data subsets.
- Ablation on TaxiBJ P1: removing region classification (w/o-MultiFC) or the whole EncM branch (w/o-STMask) raises RMSE from 15.36 to 15.80. RMSE is second-best on TaxiBJ P3(80%), P3(100%) and P4(20%), where MAE stays best.
Abstract
Intelligent Transportation Systems (ITS) are essential for modern urban development, with urban flow prediction being a key component. Accurate flow prediction optimizes routes and resource allocation, benefiting residents, businesses, and the environment. However, few methods address the spatial-temporal heterogeneity of urban flows. Existing methods typically capture spatial features solely from urban flows, but spatial feature sensitivity becomes a bottleneck when dealing with small or noisy datasets. To address this issue, we propose a method for urban flow prediction via mutual reinforcement with multi-scale regional information (MR-UFP). Firstly, we employ spatial-temporal random masking and spatial-temporal contrastive learning pre-training to directly mine spatial-temporal heterogeneity from historical flow data. Secondly, we transform the task of spatial feature extraction and embedding for urban flow prediction into a mutual reinforcement task by multi-scale region classification auxiliary task. The adaptive environment fusion module and real-time graph processing dynamically correlate regional and environmental features during flow prediction. To balance the mutual reinforcement of the two tasks, we design a joint loss function to optimize feature embedding and feedback correction, ensuring robust and accurate urban flow prediction. Extensive experiments on two real-world datasets demonstrate that MR-UFP outperforms baseline models, showcasing its robustness and effectiveness even with minimal data.
Published abstract, character-identical to PubMed 39579750; the typeset Elsevier article prints "spatial–temporal" with en dashes.
Frequently asked questions
Why is the code repository named MR-UPF when the paper says MR-UFP?
MR-UFP is the method name used throughout the Neural Networks paper, short for urban flow prediction via mutual reinforcement with multi-scale regional information. The GitHub repository github.com/CodeZx6/MR-UPF uses a different letter order; it holds the code and the hand-labeled region data, and the paper's footnote link github.com/CodeZx6/Data-of-MR-UPF now redirects to it.
Does MR-UFP help when training data is limited?
The paper trains on 20%, 40%, 60%, 80% and 100% subsets of TaxiBJ P1–P4 and BikeNYC, testing on the last 10% of each; with 20% of BikeNYC it reaches RMSE 4.15 versus 4.55 for DeepMeshCity. It is best on most subsets, but its RMSE is second-best on TaxiBJ P3(80%), P3(100%) and P4(20%), where its MAE is still lowest. The abstract motivates the method with small or noisy datasets, but the paper does not report experiments with added noise.
How are the flow-prediction and region-classification losses balanced?
The three region-classification losses (road/non-road, coarse, fine-grained) are combined with learnable weights into L_rcf. L_rcf and the prediction loss L_P are then weighted inversely to their magnitudes, giving Loss = 2·L_P·L_rcf/(L_P + L_rcf), so neither loss dominates. On TaxiBJ P1 this gives RMSE/MAE 15.36/9.16, versus 15.41/9.21 for L_P + L_rcf and 15.44/9.24 for a plain sum of all four losses.
How does MR-UFP relate to MC-STL and ST-FCL?
MC-STL (CIKM 2023) is the preliminary version of this work and uses mask-enhanced pre-training and contrastive learning with global–local cross-attention; the paper names its main new contribution as treating spatial feature extraction as a multi-scale region-classification task trained jointly with flow prediction rather than as an upstream step. ST-FCL (Knowledge-Based Systems 2023) is earlier work by the same first author on spatio-temporal fusion and contrastive learning. Both are baselines, and on full TaxiBJ MC-STL is the runner-up (RMSE 14.53 vs 14.32).
Which datasets and region labels does MR-UFP use?
MR-UFP uses TaxiBJ (Beijing taxi trajectories, 2×32×32 grid, 30-min intervals, with holiday and weather data) and BikeNYC (New York City bike trips, 2×16×8 grid, 1-h intervals, holidays but no weather data). Regions were labeled by hand from real maps at three scales, road/non-road, coarse and fine-grained, giving 2, 5 and 142 classes on TaxiBJ and 2, 8 and 24 on BikeNYC. Errors are RMSE and MAE averaged over 5 random seeds, compared against 13 baselines.
Which components matter most in the ablation?
On full TaxiBJ P1 the complete model scores RMSE 15.36 and MAE 9.16. Removing the multi-scale region-classification task (w/o-MultiFC) or the whole EncM branch (w/o-STMask) raises RMSE to 15.80; removing the contrastive EncC branch gives 15.72, removing environment features 15.66, and skipping both pre-training stages 15.52. Skipping only the masked or only the contrastive pre-training gives 15.45 or 15.47.
How do I cite MR-UFP?
Cite it as: Zhang, X., Cao, M., Gong, Y., Wu, X., Dong, X., Guo, Y., Zhao, L., & Zhang, C. (2025). Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information. Neural Networks, 182, Article 106900. https://doi.org/10.1016/j.neunet.2024.106900 BibTeX is available at https://codezx6.github.io/papers/mr-ufp.html#cite and in https://codezx6.github.io/publications.bib.
中文摘要
借助多尺度区域信息的互增强学习提升城市流量预测(MR-UFP)
现有方法通常仅从城市流量中提取空间特征,数据量小或含噪时,空间特征敏感性成为瓶颈。MR-UFP先以时空随机掩码和时空对比学习预训练,挖掘时空异质性;再借助多尺度区域分类辅助任务,经联合损失使空间特征提取与流量预测互相增强,并以自适应环境融合和实时图处理关联区域与环境特征。在TaxiBJ和BikeNYC完整数据上优于13个基线(RMSE 14.32、4.45)。
关键词:城市流量预测;互增强学习;多尺度区域信息;多尺度区域分类;时空随机掩码预训练;时空对比学习
Keywords
urban flow prediction mutual reinforcement multi-scale region information spatial–temporal systems spatial-temporal random masking contrastive pre-training multi-scale region classification intelligent transportation systems
Also referred to as: MR-UPF; urban flow prediction via mutual reinforcement with multi-scale regional information; mutual reinforcement urban flow prediction; multi-scale region classification auxiliary task; real-time graph processing module (RGPM); spatial-temporal random masking pre-training.
Research area map and related search terms
Field path (broad to narrow): artificial intelligence › machine learning › data mining › spatio-temporal data mining › time series forecasting › intelligent transportation systems › smart cities › urban computing › traffic prediction › crowd flow prediction › urban flow prediction › multi-task spatio-temporal learning › MR-UFP
Task
urban flow prediction crowd flow prediction inflow and outflow prediction traffic prediction with limited data spatio-temporal forecasting
Method
multi-task learning auxiliary region classification task multi-scale region information masked pre-training spatial-temporal contrastive pre-training environment feature fusion attention-derived graph GCN
Datasets
TaxiBJ BikeNYC
Compared with
ST-ResNet ATFM AGCRN STGCN PDFormer STAEformer DeepMeshCity MC-STL ST-FCL
中文
城市流量预测 人群流量预测 交通流量预测 多任务学习 多尺度区域 掩码预训练 对比预训练
Related areas
traffic forecasting multi-task learning auxiliary learning urban region profiling region representation learning text-enhanced spatio-temporal modeling self-supervised learning masked pre-training contrastive learning spatio-temporal graph neural networks data-efficient learning external factors (weather, holidays)
Applications
traffic management resource allocation route optimization urban planning bike-sharing operations
Related benchmarks (not used in this paper)
TaxiNYC BikeDC PEMS-BAY METR-LA PeMS04 PeMS08
中文 (扩展)
交通预测 时空数据挖掘 智能交通 智慧城市 城市计算 多任务学习 辅助任务学习 区域画像 区域表征学习 数据高效学习 天气与节假日外部因素 路径优化 资源分配
Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.
Related papers
Urban flow prediction
- ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction (Knowledge-Based Systems, vol. 282, article 111104, 2023)
- MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction (CIKM 2023, pp. 3298–3307)
Other topics
- DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech (2026)
- PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition (2026)
- DSTCN: Exploiting dynamic spatio-temporal correlations for origin-destination demand prediction (2026)
- URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics (2026)
- S2CMEN: A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image (2025)
- BiST-IF: Enhancing origin–destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution (2025)
- Automatic visual recognition for leaf disease based on enhanced attention mechanism (2024)
- S3CFSL: Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification (2024)
Cite this paper
@article{zhang2025mrufp,
title = {Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information},
author = {Zhang, Xu and Cao, Mengxin and Gong, Yongshun and Wu, Xiaoming and Dong, Xiangjun and Guo, Ying and Zhao, Long and Zhang, Chengqi},
journal = {Neural Networks},
year = {2025},
volume = {182},
pages = {106900},
doi = {10.1016/j.neunet.2024.106900},
issn = {0893-6080},
url = {https://doi.org/10.1016/j.neunet.2024.106900}
}Zhang, X., Cao, M., Gong, Y., Wu, X., Dong, X., Guo, Y., Zhao, L., & Zhang, C. (2025). Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information. Neural Networks, 182, Article 106900. https://doi.org/10.1016/j.neunet.2024.106900