MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information

Xu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang

Neural Networks, vol. 182, article 106900 (2025)

Also indexed in: Semantic Scholar · OpenAlex · dblp · PubMed · Publisher page · PolyU Institutional Research Archive

TL;DR. MR-UFP predicts the inflow and outflow of city grid regions, even with limited training data, by pre-training encoders with spatial-temporal masking and contrastive learning and training flow prediction jointly with a multi-scale region-classification task. On full TaxiBJ and BikeNYC it beats all 13 baselines (RMSE 14.32 and 4.45).

Quick facts

Method nameMR-UFP; MR-UPF (GitHub repository name, github.com/CodeZx6/MR-UPF; the paper's data footnote links github.com/CodeZx6/Data-of-MR-UPF, which now redirects there)
AuthorsXu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang
DOI10.1016/j.neunet.2024.106900
TaskUrban flow prediction: forecasting the inflow and outflow of each grid region in a city, a key component of intelligent transportation systems
Pre-trainingEncM, a ViT encoder pre-trained to reconstruct spatio-temporally masked regions (75% masking rate), and EncC, a global–local cross-attention encoder pre-trained with a triplet loss on temporally shuffled negatives
Auxiliary taskMulti-scale region classification (road/non-road, coarse, fine-grained) on hand-labeled region texts, trained jointly with flow prediction so the two tasks reinforce each other
Other modulesTime-text and Region-text Encoders, adaptive environment feature fusion, and a real-time graph processing module (RGPM) that turns EncM attention weights into an adjacency matrix for a GCN
Joint lossThree region losses combined with learnable weights into L_rcf; Loss = 2·L_P·L_rcf/(L_P + L_rcf), so the smaller loss gets more weight and neither dominates
DatasetsTaxiBJ (Beijing taxis, 2×32×32 grid, 30-min intervals, holiday and weather data) and BikeNYC (New York City bikes, 2×16×8 grid, 1-h intervals, no weather data)
Region labelsHand-labeled from real maps: 2/5/142 classes (TaxiBJ) and 2/8/24 (BikeNYC) at the three scales; released with the code at github.com/CodeZx6/MR-UPF
Metrics and baselinesRMSE and MAE averaged over 5 random seeds, with t-tests; 13 baselines including ST-ResNet, ATFM, AGCRN, STGCN, PDFormer, STAEformer, ST-FCL, MC-STL and DeepMeshCity
Headline resultFull TaxiBJ RMSE/MAE 14.32/8.56 vs 14.53/8.70 for MC-STL; full BikeNYC 4.45/2.19 vs 4.63/2.21 for DeepMeshCity
AuthorshipXu Zhang and Mengxin Cao contributed equally; Xiaoming Wu and Xiangjun Dong are corresponding authors
PublicationNeural Networks, vol. 182, article 106900 (February 2025 issue); received 19 June 2024, accepted 7 November 2024, available online 16 November 2024
AccessSubscription article, not open access. The PolyU Institutional Research Archive lists an accepted manuscript under embargo until 28 February 2027
Codehttps://github.com/CodeZx6/MR-UPF
BibTeX keyzhang2025mrufp

Main results

SettingResultCompared with
TaxiBJ, full dataset (last 28 days as test set)RMSE = 14.32, MAE = 8.56
“our proposed MR-UFP model yielded a mean RMSE of 14.32 and an MAE of 8.56. This represents an enhancement in prediction accuracy by 1.47% and 1.64% over the sub-optimal model MC-STL, respectively.”
MC-STL (strongest baseline): RMSE 14.53, MAE 8.70; p < 0.01
BikeNYC, full dataset (last 10 days as test set; no weather data)RMSE = 4.45, MAE = 2.19
“In the experiment utilizing the complete BikeNYC dataset, we used the last 10 days for testing. As shown in Table 2, MR-UFP achieved an RMSE of 4.45 and an MAE of 2.19.”
DeepMeshCity (strongest baseline): RMSE 4.63, MAE 2.21; Table 2 gives p < 0.01 (RMSE) and p < 0.03 (MAE)
BikeNYC, 20% data subset (last 10% as test set)RMSE = 4.15, MAE = 2.09
“MR-UFP (Ours) 4.15 ± 0.03 2.09 ± 0.01”
DeepMeshCity: RMSE 4.55, MAE 2.13
TaxiBJ P3, 20% data subset (last 10% as test set)RMSE = 15.57, MAE = 9.95
“MR-UFP(Ours) 15.57 ± 0.04 9.95 ± 0.03”
Best baselines: ST-FCL RMSE 16.82; MC-STL MAE 10.77
TaxiBJ subsets P3(80%), P3(100%) and P4(20%)RMSE rank = second-best; MAE rank = best
“Notably, in the datasets P3(80%), P3(100%), and P4(20%), although the RMSE of MR-UFP was sub-optimal, the MAE reached optimal levels.”
—
Ablation, full TaxiBJ P1Full model RMSE = 15.36 (MAE 9.16); without region classification (w/o-MultiFC) RMSE = 15.80 (MAE 9.35)
“the RMSE loss of the variant ‘MR-UFP-w/o-MultiFC’ increases to 15.80 when this auxiliary task is entirely removed”
Without EncM branch (w/o-STMask) 15.80; without EncC branch (w/o-STCon) 15.72; without environment features 15.66; without both pre-trainings 15.52; without masked pre-training only 15.45 (RMSE)

Quoted text is verbatim from the paper.

Key points

Abstract

Intelligent Transportation Systems (ITS) are essential for modern urban development, with urban flow prediction being a key component. Accurate flow prediction optimizes routes and resource allocation, benefiting residents, businesses, and the environment. However, few methods address the spatial-temporal heterogeneity of urban flows. Existing methods typically capture spatial features solely from urban flows, but spatial feature sensitivity becomes a bottleneck when dealing with small or noisy datasets. To address this issue, we propose a method for urban flow prediction via mutual reinforcement with multi-scale regional information (MR-UFP). Firstly, we employ spatial-temporal random masking and spatial-temporal contrastive learning pre-training to directly mine spatial-temporal heterogeneity from historical flow data. Secondly, we transform the task of spatial feature extraction and embedding for urban flow prediction into a mutual reinforcement task by multi-scale region classification auxiliary task. The adaptive environment fusion module and real-time graph processing dynamically correlate regional and environmental features during flow prediction. To balance the mutual reinforcement of the two tasks, we design a joint loss function to optimize feature embedding and feedback correction, ensuring robust and accurate urban flow prediction. Extensive experiments on two real-world datasets demonstrate that MR-UFP outperforms baseline models, showcasing its robustness and effectiveness even with minimal data.

Published abstract, character-identical to PubMed 39579750; the typeset Elsevier article prints "spatial–temporal" with en dashes.

Frequently asked questions

Why is the code repository named MR-UPF when the paper says MR-UFP?

MR-UFP is the method name used throughout the Neural Networks paper, short for urban flow prediction via mutual reinforcement with multi-scale regional information. The GitHub repository github.com/CodeZx6/MR-UPF uses a different letter order; it holds the code and the hand-labeled region data, and the paper's footnote link github.com/CodeZx6/Data-of-MR-UPF now redirects to it.

Does MR-UFP help when training data is limited?

The paper trains on 20%, 40%, 60%, 80% and 100% subsets of TaxiBJ P1–P4 and BikeNYC, testing on the last 10% of each; with 20% of BikeNYC it reaches RMSE 4.15 versus 4.55 for DeepMeshCity. It is best on most subsets, but its RMSE is second-best on TaxiBJ P3(80%), P3(100%) and P4(20%), where its MAE is still lowest. The abstract motivates the method with small or noisy datasets, but the paper does not report experiments with added noise.

How are the flow-prediction and region-classification losses balanced?

The three region-classification losses (road/non-road, coarse, fine-grained) are combined with learnable weights into L_rcf. L_rcf and the prediction loss L_P are then weighted inversely to their magnitudes, giving Loss = 2·L_P·L_rcf/(L_P + L_rcf), so neither loss dominates. On TaxiBJ P1 this gives RMSE/MAE 15.36/9.16, versus 15.41/9.21 for L_P + L_rcf and 15.44/9.24 for a plain sum of all four losses.

How does MR-UFP relate to MC-STL and ST-FCL?

MC-STL (CIKM 2023) is the preliminary version of this work and uses mask-enhanced pre-training and contrastive learning with global–local cross-attention; the paper names its main new contribution as treating spatial feature extraction as a multi-scale region-classification task trained jointly with flow prediction rather than as an upstream step. ST-FCL (Knowledge-Based Systems 2023) is earlier work by the same first author on spatio-temporal fusion and contrastive learning. Both are baselines, and on full TaxiBJ MC-STL is the runner-up (RMSE 14.53 vs 14.32).

Which datasets and region labels does MR-UFP use?

MR-UFP uses TaxiBJ (Beijing taxi trajectories, 2×32×32 grid, 30-min intervals, with holiday and weather data) and BikeNYC (New York City bike trips, 2×16×8 grid, 1-h intervals, holidays but no weather data). Regions were labeled by hand from real maps at three scales, road/non-road, coarse and fine-grained, giving 2, 5 and 142 classes on TaxiBJ and 2, 8 and 24 on BikeNYC. Errors are RMSE and MAE averaged over 5 random seeds, compared against 13 baselines.

Which components matter most in the ablation?

On full TaxiBJ P1 the complete model scores RMSE 15.36 and MAE 9.16. Removing the multi-scale region-classification task (w/o-MultiFC) or the whole EncM branch (w/o-STMask) raises RMSE to 15.80; removing the contrastive EncC branch gives 15.72, removing environment features 15.66, and skipping both pre-training stages 15.52. Skipping only the masked or only the contrastive pre-training gives 15.45 or 15.47.

How do I cite MR-UFP?

Cite it as: Zhang, X., Cao, M., Gong, Y., Wu, X., Dong, X., Guo, Y., Zhao, L., & Zhang, C. (2025). Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information. Neural Networks, 182, Article 106900. https://doi.org/10.1016/j.neunet.2024.106900 BibTeX is available at https://codezx6.github.io/papers/mr-ufp.html#cite and in https://codezx6.github.io/publications.bib.

中文摘要

借助多尺度区域信息的互增强学习提升城市流量预测(MR-UFP)

现有方法通常仅从城市流量中提取空间特征,数据量小或含噪时,空间特征敏感性成为瓶颈。MR-UFP先以时空随机掩码和时空对比学习预训练,挖掘时空异质性;再借助多尺度区域分类辅助任务,经联合损失使空间特征提取与流量预测互相增强,并以自适应环境融合和实时图处理关联区域与环境特征。在TaxiBJ和BikeNYC完整数据上优于13个基线(RMSE 14.32、4.45)。

关键词:城市流量预测;互增强学习;多尺度区域信息;多尺度区域分类;时空随机掩码预训练;时空对比学习

Keywords

urban flow prediction mutual reinforcement multi-scale region information spatial–temporal systems spatial-temporal random masking contrastive pre-training multi-scale region classification intelligent transportation systems

Also referred to as: MR-UPF; urban flow prediction via mutual reinforcement with multi-scale regional information; mutual reinforcement urban flow prediction; multi-scale region classification auxiliary task; real-time graph processing module (RGPM); spatial-temporal random masking pre-training.

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › machine learning › data mining › spatio-temporal data mining › time series forecasting › intelligent transportation systems › smart cities › urban computing › traffic prediction › crowd flow prediction › urban flow prediction › multi-task spatio-temporal learning › MR-UFP

Task

urban flow prediction crowd flow prediction inflow and outflow prediction traffic prediction with limited data spatio-temporal forecasting

Method

multi-task learning auxiliary region classification task multi-scale region information masked pre-training spatial-temporal contrastive pre-training environment feature fusion attention-derived graph GCN

Datasets

TaxiBJ BikeNYC

Compared with

ST-ResNet ATFM AGCRN STGCN PDFormer STAEformer DeepMeshCity MC-STL ST-FCL

中文

城市流量预测 人群流量预测 交通流量预测 多任务学习 多尺度区域 掩码预训练 对比预训练

Related areas

traffic forecasting multi-task learning auxiliary learning urban region profiling region representation learning text-enhanced spatio-temporal modeling self-supervised learning masked pre-training contrastive learning spatio-temporal graph neural networks data-efficient learning external factors (weather, holidays)

Applications

traffic management resource allocation route optimization urban planning bike-sharing operations

Related benchmarks (not used in this paper)

TaxiNYC BikeDC PEMS-BAY METR-LA PeMS04 PeMS08

中文 (扩展)

交通预测 时空数据挖掘 智能交通 智慧城市 城市计算 多任务学习 辅助任务学习 区域画像 区域表征学习 数据高效学习 天气与节假日外部因素 路径优化 资源分配

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Urban flow prediction

Other topics

Cite this paper

@article{zhang2025mrufp,
  title        = {Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information},
  author       = {Zhang, Xu and Cao, Mengxin and Gong, Yongshun and Wu, Xiaoming and Dong, Xiangjun and Guo, Ying and Zhao, Long and Zhang, Chengqi},
  journal      = {Neural Networks},
  year         = {2025},
  volume       = {182},
  pages        = {106900},
  doi          = {10.1016/j.neunet.2024.106900},
  issn         = {0893-6080},
  url          = {https://doi.org/10.1016/j.neunet.2024.106900}
}
Zhang, X., Cao, M., Gong, Y., Wu, X., Dong, X., Guo, Y., Zhao, L., & Zhang, C. (2025). Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information. Neural Networks, 182, Article 106900. https://doi.org/10.1016/j.neunet.2024.106900